Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
9,991-word document condensed to 136 words. OpenAI · Aug 20, 2026
TL;DR
“OpenAI o3 and OpenAI o4-mini combine state-of-the-art reasoning with full tool capabilities — web browsing, Python, image and file analysis, image generation, canvas, automations, file search, and memory. These models excel at solving complex math, coding, and scientific challenges while demonstrating strong visual perception and analysis.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| SWE-Lancer | no-tools, ic_swe, dollars_earned | 86100.00 |
| SWE-Lancer | with-tools, ic_swe, dollars_earned | 76250.00 |
| Biorisk Red-Teaming Monitor Evaluation | post-mitigation, recall | 98.7% |
| SWE-Lancer | with-tools, ic_swe, accuracy | 55.0% |
| Cyber Range Evaluation | evasion, success_rate | 51.0% |
| Cyber Range Evaluation | evasion, success_rate | 51.0% |
| Cyber Range Evaluation | vulnerability_discovery_exploitation, success_rate | 34.0% |
| Cyber Range Evaluation | network_attack_simulation, success_rate | 29.0% |
Showing top 8 of 66. See full list below.
Capability claim
- “We introduce a new cybersecurity evaluation: Cyber Range.”
Mitigations
- “mitigations include post-training our reasoning models to refuse requests to identify a person based on an image, and to refuse requests for ungrounded inferences.”
- “we have deployed significant mitigations in Preparedness risk areas, and we describe those below after the Capabilities Assessment.”
Deployment scope
- “released under Version 2 of our Preparedness Framework.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 0b08b9dee641 · version dated Aug 20, 2026.
Extracted Evaluations(66 results)
Sort by:0/66 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
/ verified | coding | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ verified | coding | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ verified | coding | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SWE-Lancer/ ic_swe | other | scored | 86100.0 dollars earned | no-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
SWE-Lancer/ ic_swe | other | scored | 76250.0 dollars earned | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Biorisk Red-Teaming Monitor Evaluation | other | scored | 98.7 recall | post-mitigationmissing: shot countmissing: languagemissing: training state | self-reported |
SWE-Lancer/ ic_swe | other | scored | 55.0 accuracy | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ evasion | other | scored | 51.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ evasion | other | scored | 51.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ vulnerability_discovery_exploitation | other | scored | 34.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ network_attack_simulation | other | scored | 29.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ vulnerability_discovery_exploitation | other | scored | 29.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ network_attack_simulation | other | scored | 25.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 24.0 pass at 1 | no-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | scored | 23.0 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 18.0 pass at 1 | no-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
Cyber Range Evaluation/ easy | other | scored | 16.0 solve count | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ easy | other | scored | 14.0 solve count | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ medium | other | scored | 9.0 solve count | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ medium | other | scored | 7.0 solve count | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ sexual_exploitative | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ self_harm_instructions | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ sexual_exploitative | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ self_harm_intent | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ self_harm_intent | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ self_harm_instructions | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ self_harm_intent | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ sexual_exploitative | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ self_harm_instructions | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.8 hallucination rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.6 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Sabotage Capability Evaluation/ ai_r&d_backdoor | other | scored | 0.6 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.5 hallucination rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.5 hallucination rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.4 hallucination rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.4 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.3 hallucination rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Sabotage Capability Evaluation/ ai_r&d_backdoor | other | scored | 0.2 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.2 hallucination rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Sabotage Capability Evaluation/ ai_r&d_backdoor | other | scored | 0.1 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ hard | other | scored | 0.0 solve count | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Evaluation/ hard | other | scored | 0.0 solve count | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ multiple_choice | other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lab-Bench | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Cyber Range Online Retailer Scenario | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Privilege Escalation Scenario | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Privilege Escalation Scenario | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Privilege Escalation Scenario | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Privilege Escalation Scenario | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range Privilege Escalation Scenario | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Vision Red-Teaming ELO Evaluation | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Vision Red-Teaming ELO Evaluation | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ multiple_choice | other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ multiple_choice | other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SWE-Lancer/ swe_manager | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
OpenAI Research Engineer Interview/ coding | other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ coding | other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Vision Red-Teaming ELO Evaluation | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Vision Red-Teaming ELO Evaluation | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| safety | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |