Model Cards / OpenAI

GPT-5.3 Codex System Card

model card10,676 words·46 min read·Aug 20, 2026·Source
Version History
Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
10,676-word document condensed to 166 words. OpenAI · Aug 20, 2026
TL;DR

GPT-5.3-Codex is the most capable agentic coding model to date, combining the frontier coding performance of GPT-5.2-Codex with the reasoning and professional knowledge capabilities of GPT-5.2. This enables it to take on long-running tasks that involve research, tool use, and complex execution.

Top benchmarks
BenchmarkVariantScore
Irregular Cyberoffensive Challenge Suitewith-tools, network_attack_simulation, success_rate100.0%
Irregular Cyberoffensive Challenge Suitewith-tools, network_attack_simulation, success_rate86.0%
Irregular Cyberoffensive Challenge Suitewith-tools, vulnerability_research_and_exploitation, success_rate80.0%
Irregular Cyberoffensive Challenge Suitewith-tools, vulnerability_research_and_exploitation, success_rate75.0%
Irregular Cyberoffensive Challenge Suitewith-tools, vulnerability_research_and_exploitation, success_rate72.0%
Irregular Cyberoffensive Challenge Suitewith-tools, network_attack_simulation, success_rate68.0%
Irregular Cyberoffensive Challenge Suitewith-tools, evasion, success_rate53.0%
Irregular Cyberoffensive Challenge Suitewith-tools, evasion, success_rate53.0%

Showing top 8 of 55. See full list below.

Capability claim
  • we trained GPT-5.3-Codex to generally provide maximally helpful support on dual-use cybersecurity topics while refusing or de-escalating operational guidance for harmful actions, including areas such as malware creation, credential theft, and chained exploitation.
Safety findings
  • cannot rule out the possibility that it may be capable enough to reach the threshold.
  • cannot rule out the possibility that it is in fact Cyber High.
  • cannot rule out Cyber High–level automation of end-to-end operations or operationally relevant vulnerability discovery and exploitation.
Deployment scope
  • deployed to date, and the first that we are treating as High under the Preparedness Framework.

Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA f52661328458 · version dated Aug 20, 2026.

Extracted Evaluations(55 results)

Sort by:0/55 rows fully reproducible (0%)
BenchmarkCategoryStateScoreSetupSource
Irregular Cyberoffensive Challenge Suite/ network_attack_simulation
otherscored
100.0
success rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Cyberoffensive Challenge Suite/ network_attack_simulation
otherscored
86.0
success rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Cyberoffensive Challenge Suite/ vulnerability_research_and_exploitation
otherscored
80.0
success rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Cyberoffensive Challenge Suite/ vulnerability_research_and_exploitation
otherscored
75.0
success rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Cyberoffensive Challenge Suite/ vulnerability_research_and_exploitation
otherscored
72.0
success rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Cyberoffensive Challenge Suite/ network_attack_simulation
otherscored
68.0
success rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Cyberoffensive Challenge Suite/ evasion
otherscored
53.0
success rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Cyberoffensive Challenge Suite/ evasion
otherscored
53.0
success rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Irregular Cyberoffensive Challenge Suite/ evasion
otherscored
52.0
success rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
/ biological_weapons
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ extremism
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ biological_weapons
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ sexual_minors
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ sexual_minors
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ illicit_violent_activities
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ hate
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ illicit_violent_activities
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ extremism
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ sexual_exploitative
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ sexual_exploitative
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ self_harm
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ self_harm
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ hate
otherscored
0.9
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ illicit_non_violent_harmful_activities
otherscored
0.9
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ illicit_non_violent_harmful_activities
otherscored
0.9
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ violence
otherscored
0.9
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Destructive Action Avoidance
otherscored
0.9
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
0.9
best of 10 mean
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ violence
otherscored
0.9
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ chemical_weapons
otherscored
0.9
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ chemical_weapons
otherscored
0.9
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ abuse
otherscored
0.8
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ abuse
otherscored
0.8
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Destructive Action Avoidance
otherscored
0.8
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Destructive Action Avoidance
otherscored
0.8
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
0.8
best of 10 mean
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Destructive Action Avoidance
otherscored
0.7
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Destructive Action Avoidance
otherscored
0.7
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
0.7
best of 10 mean
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Non-Latin Script Reasoning Tokens
otherscored
0.6
percentage
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Non-Latin Script Reasoning Tokens
otherscored
0.0
percentage
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Non-Latin Script Reasoning Tokens
otherscored
0.0
percentage
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
0.0
success rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Tacit Knowledge and Troubleshooting
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
OpenAI-Proof Q&A
othermentioned
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
OpenAI-Proof Q&A
othermentioned
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Monorepo-Bench
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Monorepo-Bench
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Monorepo-Bench
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Cyber Range
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CVE-Bench
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Capture the Flag/ professional
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ open_ended
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported