Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
10,676-word document condensed to 166 words. OpenAI · Aug 20, 2026
TL;DR
“GPT-5.3-Codex is the most capable agentic coding model to date, combining the frontier coding performance of GPT-5.2-Codex with the reasoning and professional knowledge capabilities of GPT-5.2. This enables it to take on long-running tasks that involve research, tool use, and complex execution.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| Irregular Cyberoffensive Challenge Suite | with-tools, network_attack_simulation, success_rate | 100.0% |
| Irregular Cyberoffensive Challenge Suite | with-tools, network_attack_simulation, success_rate | 86.0% |
| Irregular Cyberoffensive Challenge Suite | with-tools, vulnerability_research_and_exploitation, success_rate | 80.0% |
| Irregular Cyberoffensive Challenge Suite | with-tools, vulnerability_research_and_exploitation, success_rate | 75.0% |
| Irregular Cyberoffensive Challenge Suite | with-tools, vulnerability_research_and_exploitation, success_rate | 72.0% |
| Irregular Cyberoffensive Challenge Suite | with-tools, network_attack_simulation, success_rate | 68.0% |
| Irregular Cyberoffensive Challenge Suite | with-tools, evasion, success_rate | 53.0% |
| Irregular Cyberoffensive Challenge Suite | with-tools, evasion, success_rate | 53.0% |
Showing top 8 of 55. See full list below.
Capability claim
- “we trained GPT-5.3-Codex to generally provide maximally helpful support on dual-use cybersecurity topics while refusing or de-escalating operational guidance for harmful actions, including areas such as malware creation, credential theft, and chained exploitation.”
Safety findings
- “cannot rule out the possibility that it may be capable enough to reach the threshold.”
- “cannot rule out the possibility that it is in fact Cyber High.”
- “cannot rule out Cyber High–level automation of end-to-end operations or operationally relevant vulnerability discovery and exploitation.”
Deployment scope
- “deployed to date, and the first that we are treating as High under the Preparedness Framework.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA f52661328458 · version dated Aug 20, 2026.
Extracted Evaluations(55 results)
Sort by:0/55 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
Irregular Cyberoffensive Challenge Suite/ network_attack_simulation | other | scored | 100.0 success rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Cyberoffensive Challenge Suite/ network_attack_simulation | other | scored | 86.0 success rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Cyberoffensive Challenge Suite/ vulnerability_research_and_exploitation | other | scored | 80.0 success rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Cyberoffensive Challenge Suite/ vulnerability_research_and_exploitation | other | scored | 75.0 success rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Cyberoffensive Challenge Suite/ vulnerability_research_and_exploitation | other | scored | 72.0 success rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Cyberoffensive Challenge Suite/ network_attack_simulation | other | scored | 68.0 success rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Cyberoffensive Challenge Suite/ evasion | other | scored | 53.0 success rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Cyberoffensive Challenge Suite/ evasion | other | scored | 53.0 success rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Irregular Cyberoffensive Challenge Suite/ evasion | other | scored | 52.0 success rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ biological_weapons | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ extremism | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ biological_weapons | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ sexual_minors | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ sexual_minors | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ illicit_violent_activities | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ hate | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ illicit_violent_activities | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ extremism | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ sexual_exploitative | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ sexual_exploitative | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ self_harm | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ self_harm | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ hate | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ illicit_non_violent_harmful_activities | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ illicit_non_violent_harmful_activities | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ violence | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Destructive Action Avoidance | other | scored | 0.9 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.9 best of 10 mean | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ violence | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ chemical_weapons | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ chemical_weapons | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ abuse | other | scored | 0.8 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ abuse | other | scored | 0.8 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Destructive Action Avoidance | other | scored | 0.8 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Destructive Action Avoidance | other | scored | 0.8 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.8 best of 10 mean | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Destructive Action Avoidance | other | scored | 0.7 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Destructive Action Avoidance | other | scored | 0.7 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.7 best of 10 mean | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Non-Latin Script Reasoning Tokens | other | scored | 0.6 percentage | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Non-Latin Script Reasoning Tokens | other | scored | 0.0 percentage | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Non-Latin Script Reasoning Tokens | other | scored | 0.0 percentage | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.0 success rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
Tacit Knowledge and Troubleshooting | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI-Proof Q&A | other | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI-Proof Q&A | other | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Monorepo-Bench | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Monorepo-Bench | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Monorepo-Bench | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CVE-Bench | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Capture the Flag/ professional | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ open_ended | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |