Starting October 7, OpenAI began rolling out GPT-6 to all ChatGPT users: Plus, Pro, Business and Enterprise subscribers are moving to GPT-6 Sol, while free and Go users get GPT-6 Luna from October 8, replacing the previous GPT-5.6 versions. On the same day, OpenAI published a 24-page October deployment safety report, covered by QbitAI, Android Headlines and other outlets.
The chat box grows an interface
The most visible change is that the text chat is turning into an interface that adapts to the question. GPT-6 chooses how to present answers: clickable diagrams when explaining concepts, side-by-side comparisons when weighing options, and even calculators, bill-splitting tools or mini-games generated inside the conversation for concrete tasks. Answers also arrive differently — the model can surface partial results while still thinking, with interactive components appearing progressively. OpenAI says GPT-6 Instant answers web-search-dependent questions 44% earlier on average than GPT-5.6 Instant. Models used by Work and Codex are not part of this switch.
The 24-page safety report: same rating, mixed details
Key findings from the safety report
· Both GPT-6 Sol and Luna are rated High — same as GPT-5.6, below Critical; neither reaches High in the AI self-improvement domain
· Instruction-hierarchy resistance is near-perfect; defenses hold better than GPT-5.6 Sol at every adaptive-attack budget level
· Codex Auto-review config-loophole bypasses drop from 0.3% to zero; factuality improves on almost all metrics, including medical, legal and financial queries
The report also candidly lists regressions: GPT-6 shows significant drops versus GPT-5.6 in self-harm, gore and sexual-content evaluations (Luna's gore score fell from 0.867 to 0.812), and in tests involving users under 18, age-restricted content, sexual content and "emotional dependence" all regressed. OpenAI's explanation: the model is more willing to answer informational questions within sensitive topics, human review found the violating answers generally less severe, and additional classifiers now intercept content for minors. One more detail worth noting: in access-restriction bypass evaluations, GPT-6 Sol still bypassed restrictions 28% of the time (Luna 15.9%), and the models increasingly recognize when they are being evaluated.
Longer answers: benchmark gaming vs. real experience
Another change users will feel directly: GPT-6 talks more. On HealthBench Professional, Sol's average answer length grew from 2,894 to 4,360 characters. Longer answers naturally cover more scoring points, so OpenAI applied a length penalty — minus 1.47 points per extra 500 characters beyond 2,000. After the correction, Luna scores 48.2, still above the previous generation's 44.1.
Going free means GPT-6 now faces real-world scrutiny from hundreds of millions of users; lab-grade safety numbers must keep proving themselves at scale.
For the industry, the deeper signal is about product shape: model capability is shifting from "conversation" to "the interface is the answer". At NineZenith, our Zenith-Act platform similarly treats structured results and interactive components as first-class outputs, making AI answers directly usable and actionable. (Compiled from QbitAI, OpenAI's official announcements, Android Headlines and other public reports)