Model Cards / OpenAI

GPT-5.4 Thinking System Card

model card11,900 words·52 min read·Aug 3, 2026·Source
Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
11,900-word document condensed to 138 words. OpenAI · Aug 3, 2026
TL;DR

GPT-5.4 Thinking is the latest reasoning model in the GPT-5 series, and explained in our blog. The comprehensive safety mitigation approach for this model is similar to previous models in this series, but 5.4 Thinking is the first general purpose model to have implemented mitigations for High capability in Cybersecurity.

Top benchmarks
BenchmarkVariantScore
Irregular Atomic Challenge Suiteextended-thinking, hard, resolve_rate100.0%
Irregular Atomic Challenge Suiteextended-thinking, network_attack_simulation, success_rate88.0%
Irregular Atomic Challenge Suiteextended-thinking, medium, resolve_rate82.3%
Cyber Rangewith-tools, pass_rate80.0%
Cyber Rangewith-tools, pass_rate73.3%
Irregular Atomic Challenge Suiteextended-thinking, vulnerability_research_and_exploitation, success_rate73.0%
Cyber Rangewith-tools, pass_rate53.3%
Irregular Atomic Challenge Suiteextended-thinking, evasion, success_rate48.0%

Showing top 8 of 50. See full list below.

Capability claim
  • we trained our agents to revert their own changes after long rollouts while protecting implicit, simulated user work.
Safety findings
  • cannot rule out the possibility that it is in fact Cyber High.
Deployment scope
  • available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate.

Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 8a6edb828596 · version dated Aug 3, 2026.

Extracted Evaluations(50 results)

Sort by:0/50 rows fully reproducible (0%)
BenchmarkCategoryStateScoreSetupSource
codingcited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ verified
codingcited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ pro
knowledgecited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Irregular Atomic Challenge Suite/ hard
otherscored
100.0
resolve rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Atomic Challenge Suite/ network_attack_simulation
otherscored
88.0
success rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Atomic Challenge Suite/ medium
otherscored
82.3
resolve rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
Cyber Range
otherscored
80.0
pass rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Cyber Range
otherscored
73.3
pass rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Atomic Challenge Suite/ vulnerability_research_and_exploitation
otherscored
73.0
success rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
Cyber Range
otherscored
53.3
pass rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Atomic Challenge Suite/ evasion
otherscored
48.0
success rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
Cyber Range
otherscored
47.0
pass rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
otherscored
45.5
resolve rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
otherscored
11.0
success rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
otherscored
9.1
resolve rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
OpenAI-Proof Q&A
otherscored
8.3
pass at 1
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
/ incentivized
otherscored
6.0
accuracy drop
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Covert Deception Rate/ no_nudge
otherscored
1.0
deception rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ 10k_characters
otherscored
0.3
success rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ 10k_characters
otherscored
0.2
success rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CoT Monitorability
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CoT Monitorability
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CoT Monitorability/ impossible_tasks
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CoT Monitorability/ agentic_misalignment
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CoT Monitorability/ anti_scheming
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CoT Monitorability/ sabotage
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CoT Monitorability/ shadearena
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CoT Monitorability/ memory
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CoT Monitorability
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Multi-select Multimodal Troubleshooting Virology
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ open_ended
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Tacit Knowledge and Troubleshooting
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Tacit Knowledge and Troubleshooting
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Tacit Knowledge and Troubleshooting
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Capture the Flag/ professional
othermentioned
pass at 12
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
CVE-Bench
othermentioned
pass at 1
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
CVE-Bench
othermentioned
pass at 1
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
CVE-Bench
othermentioned
pass at 1
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Monorepo-Bench
othermentioned
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
othermentioned
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
othermentioned
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
OpenAI-Proof Q&A
othermentioned
pass at 1
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Covert Deception Rate/ no_nudge
othermentioned
deception rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Covert Deception Rate/ no_nudge
othermentioned
deception rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
reasoningcited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported