“GPT-5.5 is a new model designed for complex, real-world work, including writing code, researching online, analyzing information, creating documents and spreadsheets, and moving across tools to get things done. Relative to earlier models, GPT-5.5 understands the task earlier, asks for less guidance, uses tools more effectively, checks it work and keeps going until it’s done.”
| Benchmark | Variant | Score |
|---|---|---|
| Sandbagging QA | accuracy | 100.0% |
| Strategic Deception Capability Sandbagging | accuracy | 99.6% |
| Irregular Atomic Challenge Suite | network_attack_simulation, success_rate | 98.0% |
| Irregular Atomic Challenge Suite | vulnerability_research_and_exploitation, success_rate | 92.0% |
| UK AISI narrow cyber tasks | expert_level, pass_at_5 | 90.5% |
| UK AISI narrow cyber tasks | expert_level, pass_at_5 | 71.4% |
| UK AISI narrow cyber tasks | expert_level, pass_at_1 | 66.7% |
| Irregular Atomic Challenge Suite | evasion, success_rate | 54.0% |
Showing top 8 of 62. See full list below.
- “We are releasing GPT-5.5 with our strongest set of safeguards to date, designed to reduce misuse while preserving legitimate, beneficial uses of advanced capabilities.”
- “we have deployed an expanded set of safeguards to restrict the ability of malicious actors to benefit from increased capabilities in cybersecurity performance (section link to Cyber Safeguards section).”
- “we trained GPT-5.5 to refuse requests that clearly enable unauthorized, destructive, or harmful actions, including areas such as malware deployment, credential theft, and exfiltration.”
- “available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 0627c42bc9da · version dated Aug 20, 2026.
Extracted Evaluations(62 results)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
/ verified | coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ pro | knowledge | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 100.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 99.6 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Irregular Atomic Challenge Suite/ network_attack_simulation | other | scored | 98.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Irregular Atomic Challenge Suite/ vulnerability_research_and_exploitation | other | scored | 92.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
UK AISI narrow cyber tasks/ expert_level | other | scored | 90.5 pass at 5 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
UK AISI narrow cyber tasks/ expert_level | other | scored | 71.4 pass at 5 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
UK AISI narrow cyber tasks/ expert_level | other | scored | 66.7 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Irregular Atomic Challenge Suite/ evasion | other | scored | 54.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ sandbagging_condition | other | scored | 52.0 verbalized awareness rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 29.0 lie rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 26.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Hallucination Evaluation (Flagged Conversations) | other | scored | 23.0 claim level accuracy improvement | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Verbalized Evaluation Awareness | other | scored | 22.1 moderate or higher rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Verbalized Evaluation Awareness | other | scored | 17.3 moderate or higher rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 16.5 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 13.8 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 12.8 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Verbalized Evaluation Awareness | other | scored | 11.7 moderate or higher rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 10.0 lie rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 9.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 7.0 lie rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
OpenAI-Proof | other | scored | 5.8 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 3.5 pass at 4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Hallucination Evaluation (Flagged Conversations) | other | scored | 3.0 response level error rate reduction | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 1.5 pass at 4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Dynamic Benchmarks/ emotional_reliance | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ mental_health | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ mental_health | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ emotional_reliance | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ violent_illicit_behavior | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ self_harm | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ violent_illicit_behavior | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ mental_health | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ violent_illicit_behavior | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ self_harm | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ violent_illicit_behavior | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ emotional_reliance | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ self_harm | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ self_harm | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ emotional_reliance | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dynamic Benchmarks/ mental_health | other | scored | 0.8 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Sabotage Capability Evaluation | other | scored | 0.7 mean score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Sabotage Capability Evaluation | other | scored | 0.6 mean score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Sabotage Capability Evaluation | other | scored | 0.6 mean score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 50k_characters | other | scored | 0.5 controllability rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 50k_characters | other | scored | 0.3 controllability rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 50k_characters | other | scored | 0.2 controllability rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
First-Person Fairness | other | scored | 0.0 harm overall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
First-Person Fairness | other | scored | 0.0 harm overall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
First-Person Fairness | other | scored | 0.0 harm overall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
First-Person Fairness | other | scored | 0.0 harm overall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ non_sandbagging_condition | other | scored | 0.0 verbalized awareness rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.0 pass at 4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
CAISI Cyber Assessment | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |