“GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reasoning model for harder problems, and a real-time router that quickly decides which model to use based on conversation type, complexity, tool needs, and explicit intent (for example, if you say “think hard about this” in the prompt). The router is continuously trained on real signals, including when user”
| Benchmark | Variant | Score |
|---|---|---|
| SWE-bench | verified, pass_at_1 | 74.9% |
| Biological risk content evaluation | adversarial_production_prompts, not_unsafe | 96.8% |
| Topical classifier | recall | 96.0% |
| Biological risk content evaluation | challenging_prompts_biosafety_phds, not_unsafe | 93.6% |
| Reasoning monitor | challenging_prompts, recall | 83.8% |
| Topical classifier | f1 | 83.4% |
| Topical classifier | precision | 73.7% |
| Reasoning monitor | challenging_prompts, f1 | 73.0% |
Showing top 8 of 38. See full list below.
- “We trained gpt-5-thinking and gpt-5-thinking-mini to follow OpenAI’s safety policies, using the taxonomy of biorisk information described above and in the ChatGPT agent System Card.”
- “We believe this risk to be sufficiently minimized under our Preparedness Framework primarily because discovering such jailbreaks will be challenging, users and developers who attempt to get biorisk content may get banned (and may get reported to law enforcement in extreme cases), and because we expect to be able to discover and respond to publicly discovered jailbreaks via our bug bounty and rapid remediation programs.”
- “we believe that over-refusing on benign queries is a more likely possibility. Incrementally Leaking Higher Risk Content: This threat model considers if users may be able to incrementally ask for information that is increasingly more detailed or combine individually benign information across sessions which in totality lead to higher risk content.”
- “We believe that this risk is minimal, given the strict access conditions and our vetting processes which include assessing biosafety and security controls. Risks in the API: We have two classes of actors in the API: developers, and their end users.”
- “available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 2a743d06debd · version dated Aug 20, 2026.
Extracted Evaluations(38 results)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
/ verified | coding | scored | 74.9% pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ verified | coding | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ verified | coding | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ verified | coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| medical | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Biological risk content evaluation/ adversarial_production_prompts | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 1.0 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Biological risk content evaluation/ challenging_prompts_biosafety_phds | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ challenging_prompts | other | scored | 0.8 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.8 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.7 precision | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ challenging_prompts | other | scored | 0.7 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ challenging_prompts | other | scored | 0.6 precision | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ open_ended | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ open_ended | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
SecureBio external evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Capture the Flag | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Pattern Labs external evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
OpenAI-Proof Q&A | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
METR external evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Sandbagging evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lab-Bench | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
SWE-Lancer/ diamond_ic_swe | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| safety | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| safety | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |