Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
12,230-word document condensed to 129 words. OpenAI · Aug 20, 2026
TL;DR
“The OpenAI o model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the safety and robustness of our models.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| OpenAI Research Engineer Interview | coding, pass_at_1 | 92.0% |
| MakeMeSay | pre-mitigation, win_rate | 73.0% |
| SWE-bench | with-tools, verified, pass_at_1 | 61.0% |
| SWE-bench | no-tools, verified, pass_at_1 | 48.0% |
| SWE-bench | no-tools, verified, pass_at_1 | 39.0% |
Showing top 5 of 30. See full list below.
Capability claim
- “we have trained GPT-4o to adhere to an Instruction Hierarchy; the results for GPT-4o are for the most up-to-date model.”
Safety findings
- “not release in products) are denoted as “pre-mitigation,” specifically o3-mini (Pre-Mitigation).”
Deployment scope
- “available to us, we believe the post-mitigation o3-mini model cannot meaningfully assist in the development of radiological or nuclear weapons, but note again that this assessment is limited by what we can test.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 014bfa690d09 · version dated Aug 20, 2026.
Extracted Evaluations(30 results)
Sort by:⚠ 1 conflicting report0/30 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
/ verified | coding | scored | 61.0% pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified⚠ 1 other disagree | coding | scored | 48.0% pass at 1 | no-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified | coding | scored | 39.0% pass at 1 | no-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified | coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
OpenAI Research Engineer Interview/ coding | other | scored | 92.0 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 73.0 win rate | pre-mitigationmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Instruction Hierarchy | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Jailbreak Evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CSAW | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Lab-Bench | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
MakeMePay | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ multiple_choice | other | mentioned | — | majority-votingmissing: shot countmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ multiple_choice | other | mentioned | — | majority-votingmissing: shot countmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Agentic Tasks | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
JailbreakBench | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| safety | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| safety | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| safety | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |