Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
11,664-word document condensed to 156 words. OpenAI · Aug 20, 2026
TL;DR
“GPT-4o[1] is an autoregressive omni model, which accepts as input any combination of text, audio, image, and video and generates any combination of text, audio, and image outputs. It’s trained end-to-end across text, vision, and audio, meaning that all inputs and outputs are processed by the same neural network.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| Unauthorized Voice Generation Detection | detection_rate | 100.0% |
| OpenAI Research Coding Interview | pass_at_100 | 95.0% |
| Persuasion | audio_clips, relative_effect_size | 78.0% |
| ARC | Hausa, easy, accuracy | 71.4% |
| Persuasion | conversations, relative_effect_size | 65.0% |
| OpenAI Interview Multiple Choice Questions | cons_at_32 | 61.0% |
| Uhura-Eval | Hausa, accuracy | 59.4% |
| TruthfulQA | Yoruba, accuracy | 51.1% |
Showing top 8 of 44. See full list below.
Capability claim
- “We trained the model to adhere to behavior that would reduce risk via post-training methods and also integrated classifiers for blocking specific generations as a part of the deployed system.”
Safety findings
- “not deploy the model until mitigations lower the score to medium.”
Mitigations
- “We trained GPT-4o to refuse requests for copyrighted content, including audio, consistent with our broader practices.”
Limitations the lab flags
- “future work is needed to test whether text-audio transfer, which occurred for refusal behavior, extends to these evaluations.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 0c1b132e366b · version dated Aug 20, 2026.
Extracted Evaluations(44 results)
Sort by:⚠ 5 conflicting reports0/44 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
/ verified | coding | scored | 19.0% pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| general_knowledge | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ college_biology⚠ 16 others disagree | knowledge | scored | 0.9% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ college_biology | knowledge | scored | 0.9% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ college_biology⚠ 16 others disagree | knowledge | scored | 0.9% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ college_biology | knowledge | scored | 0.9% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ college_medicine⚠ 16 others disagree | knowledge | scored | 0.9% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ college_medicine⚠ 16 others disagree | knowledge | scored | 0.8% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ college_medicine | knowledge | scored | 0.8% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ college_medicine | knowledge | scored | 0.7% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
⚠ 16 others disagree | knowledge | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Unauthorized Voice Generation Detection | other | scored | 100.0 detection rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 95.0 pass at 100 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Persuasion/ audio_clips | other | scored | 78.0 relative effect size | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ARC/ easy | other | scored | 71.4 accuracy | Hausamissing: shot countmissing: methodmissing: training state | self-reported |
Persuasion/ conversations | other | scored | 65.0 relative effect size | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
OpenAI Interview Multiple Choice Questions | other | scored | 61.0 cons at 32 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Uhura-Eval | other | scored | 59.4 accuracy | Hausamissing: shot countmissing: methodmissing: training state | self-reported |
Uhura-Eval | other | scored | 32.3 accuracy | Hausamissing: shot countmissing: methodmissing: training state | self-reported |
ARC/ easy | other | scored | 6.1 accuracy | Hausamissing: shot countmissing: methodmissing: training state | self-reported |
Voice Output Classifier Performance | other | scored | 1.0 recall | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Voice Output Classifier Performance | other | scored | 1.0 recall | Non-Englishmissing: shot countmissing: methodmissing: training state | self-reported |
Speaker Identification Safe Behavior/ should_refuse | other | scored | 1.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Voice Output Classifier Performance | other | scored | 1.0 precision | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Voice Output Classifier Performance | other | scored | 0.9 precision | Non-Englishmissing: shot countmissing: methodmissing: training state | self-reported |
Speaker Identification Safe Behavior/ should_refuse | other | scored | 0.8 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Speaker Identification Safe Behavior/ should_comply | other | scored | 0.8 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Persuasion/ conversations_1week_followup | other | scored | 0.8 effect size | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MedMCQA/ dev | other | scored | 0.8 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
MedMCQA/ dev | other | scored | 0.8 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
MedMCQA/ dev | other | scored | 0.7 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
MedMCQA/ dev | other | scored | 0.7 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
Speaker Identification Safe Behavior/ should_comply | other | scored | 0.7 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ARA | other | scored | 0.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
METR ML Engineering Tasks | other | scored | 0.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Persuasion/ audio_clips_1week_followup | other | scored | -0.7 effect size | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Preparedness Framework Model Autonomy | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Sensitive Trait Attribution | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lambada | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Voice Safety Behavior Consistency | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
METR Long-Horizon Multi-Step Tasks | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| safety | scored | 51.1% accuracy | Yorubamissing: shot countmissing: methodmissing: training state | self-reported | |
| safety | scored | 28.3% accuracy | Yorubamissing: shot countmissing: methodmissing: training state | self-reported |