Model Cards / OpenAI

GPT-4o System Card

model card11,664 words·51 min read·Aug 20, 2026·Source
Version History
Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
11,664-word document condensed to 156 words. OpenAI · Aug 20, 2026
TL;DR

GPT-4o[1] is an autoregressive omni model, which accepts as input any combination of text, audio, image, and video and generates any combination of text, audio, and image outputs. It’s trained end-to-end across text, vision, and audio, meaning that all inputs and outputs are processed by the same neural network.

Top benchmarks
BenchmarkVariantScore
Unauthorized Voice Generation Detectiondetection_rate100.0%
OpenAI Research Coding Interviewpass_at_10095.0%
Persuasionaudio_clips, relative_effect_size78.0%
ARCHausa, easy, accuracy71.4%
Persuasionconversations, relative_effect_size65.0%
OpenAI Interview Multiple Choice Questionscons_at_3261.0%
Uhura-EvalHausa, accuracy59.4%
TruthfulQAYoruba, accuracy51.1%

Showing top 8 of 44. See full list below.

Capability claim
  • We trained the model to adhere to behavior that would reduce risk via post-training methods and also integrated classifiers for blocking specific generations as a part of the deployed system.
Safety findings
  • not deploy the model until mitigations lower the score to medium.
Mitigations
  • We trained GPT-4o to refuse requests for copyrighted content, including audio, consistent with our broader practices.
Limitations the lab flags
  • future work is needed to test whether text-audio transfer, which occurred for refusal behavior, extends to these evaluations.

Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 0c1b132e366b · version dated Aug 20, 2026.

Extracted Evaluations(44 results)

Sort by:5 conflicting reports0/44 rows fully reproducible (0%)
BenchmarkCategoryStateScoreSetupSource
/ verified
codingscored
19.0%
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
general_knowledgementioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ college_biology16 others disagree
knowledgescored
0.9%
accuracy
0-shotmissing: methodmissing: languagemissing: training state
self-reported
/ college_biology
knowledgescored
0.9%
accuracy
5-shotmissing: methodmissing: languagemissing: training state
self-reported
/ college_biology16 others disagree
knowledgescored
0.9%
accuracy
5-shotmissing: methodmissing: languagemissing: training state
self-reported
/ college_biology
knowledgescored
0.9%
accuracy
0-shotmissing: methodmissing: languagemissing: training state
self-reported
/ college_medicine16 others disagree
knowledgescored
0.9%
accuracy
5-shotmissing: methodmissing: languagemissing: training state
self-reported
/ college_medicine16 others disagree
knowledgescored
0.8%
accuracy
0-shotmissing: methodmissing: languagemissing: training state
self-reported
/ college_medicine
knowledgescored
0.8%
accuracy
5-shotmissing: methodmissing: languagemissing: training state
self-reported
/ college_medicine
knowledgescored
0.7%
accuracy
0-shotmissing: methodmissing: languagemissing: training state
self-reported
16 others disagree
knowledgementioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Unauthorized Voice Generation Detection
otherscored
100.0
detection rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
95.0
pass at 100
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Persuasion/ audio_clips
otherscored
78.0
relative effect size
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
ARC/ easy
otherscored
71.4
accuracy
Hausamissing: shot countmissing: methodmissing: training state
self-reported
Persuasion/ conversations
otherscored
65.0
relative effect size
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
OpenAI Interview Multiple Choice Questions
otherscored
61.0
cons at 32
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Uhura-Eval
otherscored
59.4
accuracy
Hausamissing: shot countmissing: methodmissing: training state
self-reported
Uhura-Eval
otherscored
32.3
accuracy
Hausamissing: shot countmissing: methodmissing: training state
self-reported
ARC/ easy
otherscored
6.1
accuracy
Hausamissing: shot countmissing: methodmissing: training state
self-reported
Voice Output Classifier Performance
otherscored
1.0
recall
ENmissing: shot countmissing: methodmissing: training state
self-reported
Voice Output Classifier Performance
otherscored
1.0
recall
Non-Englishmissing: shot countmissing: methodmissing: training state
self-reported
Speaker Identification Safe Behavior/ should_refuse
otherscored
1.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Voice Output Classifier Performance
otherscored
1.0
precision
ENmissing: shot countmissing: methodmissing: training state
self-reported
Voice Output Classifier Performance
otherscored
0.9
precision
Non-Englishmissing: shot countmissing: methodmissing: training state
self-reported
Speaker Identification Safe Behavior/ should_refuse
otherscored
0.8
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Speaker Identification Safe Behavior/ should_comply
otherscored
0.8
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Persuasion/ conversations_1week_followup
otherscored
0.8
effect size
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
MedMCQA/ dev
otherscored
0.8
accuracy
5-shotmissing: methodmissing: languagemissing: training state
self-reported
MedMCQA/ dev
otherscored
0.8
accuracy
0-shotmissing: methodmissing: languagemissing: training state
self-reported
MedMCQA/ dev
otherscored
0.7
accuracy
5-shotmissing: methodmissing: languagemissing: training state
self-reported
MedMCQA/ dev
otherscored
0.7
accuracy
0-shotmissing: methodmissing: languagemissing: training state
self-reported
Speaker Identification Safe Behavior/ should_comply
otherscored
0.7
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
ARA
otherscored
0.0
success rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
METR ML Engineering Tasks
otherscored
0.0
success rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Persuasion/ audio_clips_1week_followup
otherscored
-0.7
effect size
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Preparedness Framework Model Autonomy
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Sensitive Trait Attribution
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Lambada
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Voice Safety Behavior Consistency
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
METR Long-Horizon Multi-Step Tasks
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
reasoningmentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
safetyscored
51.1%
accuracy
Yorubamissing: shot countmissing: methodmissing: training state
self-reported
safetyscored
28.3%
accuracy
Yorubamissing: shot countmissing: methodmissing: training state
self-reported