Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
9,902-word document condensed to 126 words. OpenAI · Aug 20, 2026
TL;DR
“We’re releasing a research preview of OpenAI GPT-4.5, our largest and most knowledgeable model yet. Building on GPT-4o, GPT-4.5 scales pre-training further and is designed to be more general-purpose than our powerful STEM-focused reasoning models.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| SWE-Lancer | post-mitigation, diamond, dollars_earned | 41625.00 |
| OpenAI Research Engineer Interview | pre-mitigation, multiple_choice, accuracy | 80.0% |
| OpenAI Research Engineer Interview | multiple_choice, accuracy | 80.0% |
| OpenAI Research Engineer Interview | multiple_choice, accuracy | 80.0% |
| OpenAI Research Engineer Interview | post-mitigation, multiple_choice, accuracy | 80.0% |
| OpenAI Research Engineer Interview | coding, accuracy | 79.0% |
| OpenAI Research Engineer Interview | coding, accuracy | 79.0% |
| SWE-Lancer | diamond, pass_at_1 | 46.0% |
Showing top 8 of 90. See full list below.
Capability claim
- “We trained it using new supervision techniques combined with traditional methods like supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF), similar to those used for GPT-4o.”
Deployment scope
- “available to us, we believe that GPT-4.5 cannot meaningfully assist in the development of radiological or nuclear weapons, but note again that this assessment is limited by what we can test.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA a9f517c3cc29 · version dated Aug 20, 2026.
Extracted Evaluations(90 results)
Sort by:⚠ 15 conflicting reports0/90 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
/ verified | coding | scored | 38.0% pass at 1 | post-mitigationmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified | coding | scored | 35.0% pass at 1 | pre-mitigationmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified | coding | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ verified | coding | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| knowledge | scored | 0.9% accuracy | 0-shotENmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotSpanishmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotItalianmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotENmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotPortuguese (Brazil)missing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotFrenchmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotGermanmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotArabicmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotChinese (Simplified)missing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotJapanesemissing: methodmissing: training state | self-reported | |
⚠ 16 others disagree | knowledge | scored | 0.9% accuracy | 0-shotENmissing: methodmissing: training state | self-reported |
| knowledge | scored | 0.9% accuracy | 0-shotIndonesianmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotSpanishmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotHindimissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotKoreanmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotPortuguese (Brazil)missing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotFrenchmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotItalianmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotBengalimissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotIndonesianmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotChinese (Simplified)missing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotJapanesemissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotKoreanmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotArabicmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotHindimissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotSwahilimissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotGermanmissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.8% accuracy | 0-shotBengalimissing: methodmissing: training state | self-reported | |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotFrenchmissing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotItalianmissing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotSpanishmissing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotChinese (Simplified)missing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotIndonesianmissing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotGermanmissing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotPortuguese (Brazil)missing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotJapanesemissing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotArabicmissing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotKoreanmissing: methodmissing: training state | self-reported |
| knowledge | scored | 0.8% accuracy | 0-shotSwahilimissing: methodmissing: training state | self-reported | |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotHindimissing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotBengalimissing: methodmissing: training state | self-reported |
⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotSwahilimissing: methodmissing: training state | self-reported |
| knowledge | scored | 0.8% accuracy | 0-shotYorubamissing: methodmissing: training state | self-reported | |
| knowledge | scored | 0.7% accuracy | 0-shotYorubamissing: methodmissing: training state | self-reported | |
⚠ 16 others disagree | knowledge | scored | 0.6% accuracy | 0-shotYorubamissing: methodmissing: training state | self-reported |
SWE-Lancer/ diamond | other | scored | 41625.0 dollars earned | post-mitigationmissing: shot countmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ multiple_choice | other | scored | 80.0 accuracy | pre-mitigationmissing: shot countmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ multiple_choice | other | scored | 80.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ multiple_choice | other | scored | 80.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ multiple_choice | other | scored | 80.0 accuracy | post-mitigationmissing: shot countmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ coding | other | scored | 79.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI Research Engineer Interview/ coding | other | scored | 79.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SWE-Lancer/ diamond | other | scored | 46.0 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
METR Time Horizon | other | scored | 30.0 time horizon minutes | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SWE-Lancer/ diamond | other | scored | 20.0 pass at 1 | post-mitigationmissing: shot countmissing: languagemissing: training state | self-reported |
Scheming Reasoning Evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Wildchat/ toxic | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Wildchat/ toxic | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Wildchat/ toxic | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Scheming Reasoning Evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Scheming Reasoning Evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
METR General Autonomy and AI R&D Tasks | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Agentic Tasks | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
OpenAI Research Engineer Interview/ coding | other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SWE-Lancer/ diamond | other | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SWE-Lancer/ diamond | other | mentioned | — dollars earned | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
MakeMePay | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lab-Bench | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Biological Threat Creation Early Warning System | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
JailbreakBench | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
PAIR | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
DAN | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| safety | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| safety | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| safety | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |