Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
11,900-word document condensed to 138 words. OpenAI · Aug 3, 2026
TL;DR
“GPT-5.4 Thinking is the latest reasoning model in the GPT-5 series, and explained in our blog. The comprehensive safety mitigation approach for this model is similar to previous models in this series, but 5.4 Thinking is the first general purpose model to have implemented mitigations for High capability in Cybersecurity.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| Irregular Atomic Challenge Suite | extended-thinking, hard, resolve_rate | 100.0% |
| Irregular Atomic Challenge Suite | extended-thinking, network_attack_simulation, success_rate | 88.0% |
| Irregular Atomic Challenge Suite | extended-thinking, medium, resolve_rate | 82.3% |
| Cyber Range | with-tools, pass_rate | 80.0% |
| Cyber Range | with-tools, pass_rate | 73.3% |
| Irregular Atomic Challenge Suite | extended-thinking, vulnerability_research_and_exploitation, success_rate | 73.0% |
| Cyber Range | with-tools, pass_rate | 53.3% |
| Irregular Atomic Challenge Suite | extended-thinking, evasion, success_rate | 48.0% |
Showing top 8 of 50. See full list below.
Capability claim
- “we trained our agents to revert their own changes after long rollouts while protecting implicit, simulated user work.”
Safety findings
- “cannot rule out the possibility that it is in fact Cyber High.”
Deployment scope
- “available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 8a6edb828596 · version dated Aug 3, 2026.
Extracted Evaluations(50 results)
Sort by:0/50 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ verified | coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ pro | knowledge | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Irregular Atomic Challenge Suite/ hard | other | scored | 100.0 resolve rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Atomic Challenge Suite/ network_attack_simulation | other | scored | 88.0 success rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Atomic Challenge Suite/ medium | other | scored | 82.3 resolve rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
Cyber Range | other | scored | 80.0 pass rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Cyber Range | other | scored | 73.3 pass rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Atomic Challenge Suite/ vulnerability_research_and_exploitation | other | scored | 73.0 success rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
Cyber Range | other | scored | 53.3 pass rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Atomic Challenge Suite/ evasion | other | scored | 48.0 success rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
Cyber Range | other | scored | 47.0 pass rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
| other | scored | 45.5 resolve rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | scored | 11.0 success rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | scored | 9.1 resolve rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
OpenAI-Proof Q&A | other | scored | 8.3 pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ incentivized | other | scored | 6.0 accuracy drop | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Covert Deception Rate/ no_nudge | other | scored | 1.0 deception rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 10k_characters | other | scored | 0.3 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 10k_characters | other | scored | 0.2 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CoT Monitorability | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CoT Monitorability | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CoT Monitorability/ impossible_tasks | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CoT Monitorability/ agentic_misalignment | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CoT Monitorability/ anti_scheming | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CoT Monitorability/ sabotage | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CoT Monitorability/ shadearena | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CoT Monitorability/ memory | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
CoT Monitorability | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Multi-select Multimodal Troubleshooting Virology | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ open_ended | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Capture the Flag/ professional | other | mentioned | — pass at 12 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
CVE-Bench | other | mentioned | — pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
CVE-Bench | other | mentioned | — pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
CVE-Bench | other | mentioned | — pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Monorepo-Bench | other | mentioned | — | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
| other | mentioned | — | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
OpenAI-Proof Q&A | other | mentioned | — pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Covert Deception Rate/ no_nudge | other | mentioned | — deception rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Covert Deception Rate/ no_nudge | other | mentioned | — deception rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| reasoning | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |