Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
21,019-word document condensed to 150 words. OpenAI · Aug 3, 2026
TL;DR
“GPT-5.6 is a new family of three models: Sol, our new flagship model; Terra, a capable lower-cost option; and Luna, our fastest and most cost-efficient model. The safeguards we have built for this launch—our most robust yet— are built to deliver these models safely and at scale, around the world.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| Tacit Knowledge and Troubleshooting | mcq, accuracy_with_refusal_adjustment | 84.1% |
| Tacit Knowledge and Troubleshooting | mcq, accuracy_with_refusal_adjustment | 83.8% |
| Tacit Knowledge and Troubleshooting | mcq, accuracy | 65.0% |
| Multimodal Troubleshooting Virology | accuracy | 55.5% |
| TroubleshootingBench | accuracy | 48.0% |
| ProtocolQA | open_ended, accuracy | 43.5% |
| DNA Sequence Design for Transcription Factor Binding | pass_at_1 | 16.5% |
| DNA Sequence Design for Transcription Factor Binding | pass_at_1 | 13.8% |
Showing top 8 of 50. See full list below.
Capability claim
- “we trained the models to maintain a strong standard of overwrite avoidance while improving autonomy without relying on extra cautious prompting.”
Mitigations
- “We have deployed an expanded set of safeguards to restrict the ability of malicious actors to benefit from increased capabilities in cybersecurity performance.”
- “we have deployed Preparedness Safeguards.”
Deployment scope
- “available to the public, we can continue to reserve the most sensitive cybersecurity and biological capabilities for trusted defenders.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA a273db38c41c · version dated Aug 3, 2026.
Extracted Evaluations(50 results)
Sort by:0/50 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ verified | coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ pro | knowledge | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting/ mcq | other | scored | 84.1 accuracy with refusal adjustment | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting/ mcq | other | scored | 83.8 accuracy with refusal adjustment | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting/ mcq | other | scored | 65.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 55.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 48.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ open_ended | other | scored | 43.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 16.5 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 13.8 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 13.7 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 12.8 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 7.6 pass at 4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 3.5 pass at 4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ 5k_tokens | other | scored | 1.3 control rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 5k_tokens | other | scored | 0.7 control rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AAV_Capsid_Packaging_Prediction | other | scored | 0.5 spearman correlation | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AAV_Capsid_Packaging_Prediction | other | scored | 0.5 spearman correlation | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.4 pass at 4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ 5k_tokens | other | scored | 0.4 control rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AAV_Capsid_Packaging_Prediction | other | scored | 0.3 spearman correlation | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.0 pass at 4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Scruples | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Scruples/ suggest_right | other | mentioned | — g mean 2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Scruples/ first_person | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Flaky Tools | other | mentioned | — g mean 2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Impossible Coding Tasks | other | mentioned | — g mean 2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Honesty | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Model Spec | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Instruction Hierarchy | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Deployment Simulation | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
First-Person Fairness | other | mentioned | — harm overall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ open_ended | other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ open_ended | other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting/ mcq | other | mentioned | — accuracy with refusal adjustment | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting/ mcq | other | mentioned | — accuracy with refusal adjustment | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting/ mcq | other | mentioned | — accuracy with refusal adjustment | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Agentic Misalignment | other | mentioned | — g mean 2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Agentic Misalignment/ destructive_actions | other | mentioned | — g mean 2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Agentic Misalignment/ background_work | other | mentioned | — g mean 2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Health Queries/ patient_opinion | other | mentioned | — tpr | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Impossible Tasks | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| reasoning | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |