GPT-6 Sol is rolling out to eligible paid ChatGPT users, while GPT-6 Luna is rolling out to Free and Go users.
The October versions replace the earlier GPT-5.6 Sol and GPT-5.6 Luna models in the consumer product and versions used inside Codex and ChatGPT Work remain the September releases for now.
The accompanying system card is more interesting than a routine launch note. It shows models that are harder to jailbreak overall, yet register clear regressions on several individual safety benchmarks, and that tension is worth examining carefully.
What “High Capability” Actually Means
Under OpenAI’s Preparedness Framework, the company is treating both October GPT-6 models as High capability in cybersecurity and in biological and chemical domains. Neither model reaches the High threshold for AI self-improvement, and in biology and chemistry, Sol did not cross the reported indicative Critical thresholds, while OpenAI says separate Critical testing was not required for Luna because it scored below GPT-5.6 Sol on the High-capability evaluations.
Under OpenAI’s Preparedness Framework, “High capability” means a model could amplify existing pathways to severe harm. Models at this level must have safeguards that sufficiently minimize the associated risk before they are deployed.
Stronger Jailbreak Resistance, Mixed Safety Scores
Although the October models outperform GPT-5.6 Sol across the tested attacker budgets, their adaptive multi-turn jailbreak scores are slightly lower than the September GPT-6 versions, with broadly overlapping confidence intervals.
The models also show reductions in dishonesty, deception, and attempts to circumvent guardrails. Alignment evaluations that measure whether the models respect restrictions or hide limitations generally improved. OpenAI’s separate dynamic multi-turn evaluations for mental health, emotional reliance, and self-harm did not show statistically significant differences from the corresponding GPT-5.6 models.
At the same time, the Production Benchmarks that use challenging, production-derived prompts reveal statistically significant regressions relative to the corresponding GPT-5.6 August models where GPT-6 Sol showed a statistically significant regression on the standard self-harm safety benchmark, while GPT-6 Luna regressed on standard self-harm, gore and sexual-content benchmarks.
Separately, on OpenAI’s U18 safety evaluations, both models regressed significantly on age-restricted content, sexual content, and emotional reliance; Luna also regressed on gore.
Both models regressed on age-restricted content, sexual content, and emotional reliance. Luna suffered an additional regression on gore.
OpenAI says these benchmarks intentionally stress difficult long-tail cases and do not include all runtime safeguards, so the raw scores should not be treated as direct estimates of typical-user harm rates.
The manual review found the violative responses were borderline but generally low severity; for example, the models sometimes answered informational self-harm questions while still refusing requests that would facilitate self-harm.
For U18 users, OpenAI applies an additional classifier-based block to responses involving self-harm, sexual content, or gore; these protections are not reflected in the raw U18 benchmark scores.
Helpfulness Versus Refusal
One pattern that stands out is a shift on the safety-versus-helpfulness trade-off benchmarked via two key operational changes where OpenAI’s Safety Pareto analysis shows a trade-off: the October models are less likely to refuse benign requests or add excessive caveats, but they score lower on some harmful-request safety categories.
Whether that trade-off is acceptable depends on how heavily one weighs over-refusal versus under-refusal. OpenAI presents both the improvements and the regressions without claiming the overall safety profile is unambiguously stronger.
Capability Context and Remaining Safeguards
OpenAI says these High capability classifications match those already assigned to the corresponding GPT-5.6 models. The High classifications in cyber and biological/chemical capability are accompanied by the corresponding safeguards. On biology-related refusal evaluations, the October models actually improved relative to their GPT-5.6 counterparts. Cybersecurity safety scores remained broadly comparable.
Factuality evaluations on difficult prompt sets also improved in most metrics. B models improved over their GPT-5.6 counterparts on HealthBench Professional. Luna improved significantly across MentalHealthBench acuity levels, while Sol also improved, with its largest gain on emergent conversations.

Reading the Results Carefully
Benchmark regressions do not automatically translate into higher rates of harmful behavior for typical users. The evaluations target hard, long-tail cases. System-level protections—classifiers, crisis-resource routing, parental controls and other layers—sit on top of the base model.
Still, the decision to publish the regressions rather than emphasize only the gains is notable. It gives outside observers a clearer view of where the new models moved forward and where they moved sideways or backward.
The practical question for users and developers is whether the net change improves real-world outcomes. OpenAI’s own data shows a more resistant model on jailbreaks and deception, alongside measurable drops on specific safety categories that the company judges to be low-severity after review.
You must still decide whether the published gains in jailbreak resistance outweigh the measured drops on specific safety categories when judging the October models for their own use cases.