Model Cards / OpenAI

o3-mini System Card

model card12,230 words·53 min read·Aug 20, 2026·Source
Version History
Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
12,230-word document condensed to 129 words. OpenAI · Aug 20, 2026
TL;DR

The OpenAI o model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the safety and robustness of our models.

Top benchmarks
BenchmarkVariantScore
OpenAI Research Engineer Interviewcoding, pass_at_192.0%
MakeMeSaypre-mitigation, win_rate73.0%
SWE-benchwith-tools, verified, pass_at_161.0%
SWE-benchno-tools, verified, pass_at_148.0%
SWE-benchno-tools, verified, pass_at_139.0%

Showing top 5 of 30. See full list below.

Capability claim
  • we have trained GPT-4o to adhere to an Instruction Hierarchy; the results for GPT-4o are for the most up-to-date model.
Safety findings
  • not release in products) are denoted as “pre-mitigation,” specifically o3-mini (Pre-Mitigation).
Deployment scope
  • available to us, we believe the post-mitigation o3-mini model cannot meaningfully assist in the development of radiological or nuclear weapons, but note again that this assessment is limited by what we can test.

Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 014bfa690d09 · version dated Aug 20, 2026.

Extracted Evaluations(30 results)

Sort by:1 conflicting report0/30 rows fully reproducible (0%)
BenchmarkCategoryStateScoreSetupSource
/ verified
codingscored
61.0%
pass at 1
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
/ verified1 other disagree
codingscored
48.0%
pass at 1
no-toolsmissing: shot countmissing: languagemissing: training state
self-reported
/ verified
codingscored
39.0%
pass at 1
no-toolsmissing: shot countmissing: languagemissing: training state
self-reported
/ verified
codingcited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
codingcited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
knowledgecited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
OpenAI Research Engineer Interview/ coding
otherscored
92.0
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
73.0
win rate
pre-mitigationmissing: shot countmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Instruction Hierarchy
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Jailbreak Evaluations
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CSAW
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Lab-Bench
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
MakeMePay
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
OpenAI Research Engineer Interview/ multiple_choice
othermentioned
majority-votingmissing: shot countmissing: languagemissing: training state
self-reported
OpenAI Research Engineer Interview/ multiple_choice
othermentioned
majority-votingmissing: shot countmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Agentic Tasks
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
JailbreakBench
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
safetymentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
safetymentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
safetymentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported