Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
7,246-word document condensed to 142 words. OpenAI · Aug 20, 2026
TL;DR
“GPT-5.2 is the latest model family in the GPT-5 series, and explained in our blog. The comprehensive safety mitigation approach for these models is largely the same as that described in the GPT-5 System Card and GPT-5.1 System Card.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| CharXiv | without-mitigations, missing_image_strict_output, deception_rate | 88.8% |
| CharXiv | without-mitigations, missing_image_lenient_output, deception_rate | 54.0% |
| CharXiv | without-mitigations, missing_image_strict_output, deception_rate | 34.3% |
| CharXiv | without-mitigations, missing_image_lenient_output, deception_rate | 34.1% |
| Coding Deception | without-mitigations, deception_rate | 25.6% |
| Coding Deception | without-mitigations, deception_rate | 17.6% |
| Deception Evaluation | without-mitigations, adversarial, deception_rate | 11.8% |
| Browsing Broken Tools | without-mitigations, deception_rate | 9.4% |
Showing top 8 of 81. See full list below.
Capability claim
- “We trained gpt-5.2-thinking integrations to provide maximally helpful support on educational/cy- bersecurity topics while refusing or de-escalating operational guidance for cyber abuse, including areas such as malware creation, credential theft, and chained exploitation.”
Mitigations
- “we have deployed system-level safeguards in ChatGPT intended to mitigate this behavior.”
Deployment scope
- “available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA a2a0676aa905 · version dated Aug 20, 2026.
Extracted Evaluations(81 results)
Sort by:0/81 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| knowledge | scored | 0.9% accuracy | 0-shotcotSpanishmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotItalianmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotPortuguesemissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotPortuguesemissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotSpanishmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotIndonesianmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotItalianmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotIndonesianmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotGermanmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotArabicmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotChinesemissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotFrenchmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotArabicmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotChinesemissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotHindimissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotHindimissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotFrenchmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotJapanesemissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotJapanesemissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotKoreanmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotGermanmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotKoreanmissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotBengalimissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotBengalimissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotSwahilimissing: training state | self-reported | |
| knowledge | scored | 0.9% accuracy | 0-shotcotSwahilimissing: training state | self-reported | |
| knowledge | scored | 0.8% accuracy | 0-shotcotYorubamissing: training state | self-reported | |
| knowledge | scored | 0.8% accuracy | 0-shotcotYorubamissing: training state | self-reported | |
/ consensus | medical | scored | 1.0 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ consensus | medical | scored | 0.9 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ consensus | medical | scored | 0.9 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ consensus | medical | scored | 0.9 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| medical | scored | 0.6 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| medical | scored | 0.6 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| medical | scored | 0.5 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| medical | scored | 0.5 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ hard | medical | scored | 0.4 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ hard | medical | scored | 0.4 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ hard | medical | scored | 0.2 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ hard | medical | scored | 0.2 rubric score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CharXiv/ missing_image_strict_output | other | scored | 88.8 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported |
CharXiv/ missing_image_lenient_output | other | scored | 54.0 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported |
CharXiv/ missing_image_strict_output | other | scored | 34.3 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported |
CharXiv/ missing_image_lenient_output | other | scored | 34.1 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported |
| other | scored | 25.6 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | scored | 17.6 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported | |
Deception Evaluation/ adversarial | other | scored | 11.8 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported |
| other | scored | 9.4 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | scored | 9.1 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | scored | 8.0 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Deception Evaluation/ production_traffic | other | scored | 7.7 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported |
Deception Evaluation/ adversarial | other | scored | 5.4 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported |
Deception Evaluation/ production_traffic | other | scored | 1.6 deception rate | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported |
Cyber Safety Evaluation/ synthetic_data | other | scored | 1.0 policy compliance rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Safety Evaluation/ production_traffic | other | scored | 1.0 policy compliance rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ illicit | other | scored | 1.0 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Safety Evaluation/ synthetic_data | other | scored | 0.9 policy compliance rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Safety Evaluation/ synthetic_data | other | scored | 0.9 policy compliance rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Safety Evaluation/ production_traffic | other | scored | 0.9 policy compliance rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Safety Evaluation/ production_traffic | other | scored | 0.9 policy compliance rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ illicit | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ illicit | other | scored | 0.9 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ illicit | other | scored | 0.8 not unsafe | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
First-Person Fairness | other | scored | 0.0 harm overall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
First-Person Fairness | other | scored | 0.0 harm overall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Capture the Flag | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ open_ended | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Apollo Research Scheming Evaluation | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — pass at 1 | no-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | mentioned | — pass at 1 | no-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
External Evaluations for Cyber Capabilities | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber Range | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CVE-Bench | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Cybersecurity Evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |