Model Cards / OpenAI

GPT-5.5 System Card

model card14,626 words·64 min read·Aug 20, 2026·Source
Version History
Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
14,626-word document condensed to 190 words. OpenAI · Aug 20, 2026
TL;DR

GPT-5.5 is a new model designed for complex, real-world work, including writing code, researching online, analyzing information, creating documents and spreadsheets, and moving across tools to get things done. Relative to earlier models, GPT-5.5 understands the task earlier, asks for less guidance, uses tools more effectively, checks it work and keeps going until it’s done.

Top benchmarks
BenchmarkVariantScore
Sandbagging QAaccuracy100.0%
Strategic Deception Capability Sandbaggingaccuracy99.6%
Irregular Atomic Challenge Suitenetwork_attack_simulation, success_rate98.0%
Irregular Atomic Challenge Suitevulnerability_research_and_exploitation, success_rate92.0%
UK AISI narrow cyber tasksexpert_level, pass_at_590.5%
UK AISI narrow cyber tasksexpert_level, pass_at_571.4%
UK AISI narrow cyber tasksexpert_level, pass_at_166.7%
Irregular Atomic Challenge Suiteevasion, success_rate54.0%

Showing top 8 of 62. See full list below.

Capability claim
  • We are releasing GPT-5.5 with our strongest set of safeguards to date, designed to reduce misuse while preserving legitimate, beneficial uses of advanced capabilities.
Mitigations
  • we have deployed an expanded set of safeguards to restrict the ability of malicious actors to benefit from increased capabilities in cybersecurity performance (section link to Cyber Safeguards section).
  • we trained GPT-5.5 to refuse requests that clearly enable unauthorized, destructive, or harmful actions, including areas such as malware deployment, credential theft, and exfiltration.
Deployment scope
  • available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate.

Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 0627c42bc9da · version dated Aug 20, 2026.

Extracted Evaluations(62 results)

Sort by:0/62 rows fully reproducible (0%)
BenchmarkCategoryStateScoreSetupSource
/ verified
codingcited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
codingcited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ pro
knowledgecited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
100.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
99.6
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Irregular Atomic Challenge Suite/ network_attack_simulation
otherscored
98.0
success rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Irregular Atomic Challenge Suite/ vulnerability_research_and_exploitation
otherscored
92.0
success rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
UK AISI narrow cyber tasks/ expert_level
otherscored
90.5
pass at 5
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
UK AISI narrow cyber tasks/ expert_level
otherscored
71.4
pass at 5
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
UK AISI narrow cyber tasks/ expert_level
otherscored
66.7
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Irregular Atomic Challenge Suite/ evasion
otherscored
54.0
success rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ sandbagging_condition
otherscored
52.0
verbalized awareness rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
29.0
lie rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
26.0
success rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Hallucination Evaluation (Flagged Conversations)
otherscored
23.0
claim level accuracy improvement
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Verbalized Evaluation Awareness
otherscored
22.1
moderate or higher rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Verbalized Evaluation Awareness
otherscored
17.3
moderate or higher rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
16.5
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
13.8
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
12.8
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Verbalized Evaluation Awareness
otherscored
11.7
moderate or higher rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
10.0
lie rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
9.0
success rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
7.0
lie rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
OpenAI-Proof
otherscored
5.8
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
3.5
pass at 4
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Hallucination Evaluation (Flagged Conversations)
otherscored
3.0
response level error rate reduction
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
1.5
pass at 4
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ emotional_reliance
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ mental_health
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ mental_health
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ emotional_reliance
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ violent_illicit_behavior
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ self_harm
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ violent_illicit_behavior
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ mental_health
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ violent_illicit_behavior
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ self_harm
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ violent_illicit_behavior
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ emotional_reliance
otherscored
1.0
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ self_harm
otherscored
0.9
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ self_harm
otherscored
0.9
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ emotional_reliance
otherscored
0.9
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Dynamic Benchmarks/ mental_health
otherscored
0.8
not unsafe
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Sabotage Capability Evaluation
otherscored
0.7
mean score
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Sabotage Capability Evaluation
otherscored
0.6
mean score
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Sabotage Capability Evaluation
otherscored
0.6
mean score
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ 50k_characters
otherscored
0.5
controllability rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ 50k_characters
otherscored
0.3
controllability rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ 50k_characters
otherscored
0.2
controllability rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
First-Person Fairness
otherscored
0.0
harm overall
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
First-Person Fairness
otherscored
0.0
harm overall
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
First-Person Fairness
otherscored
0.0
harm overall
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
First-Person Fairness
otherscored
0.0
harm overall
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ non_sandbagging_condition
otherscored
0.0
verbalized awareness rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
0.0
pass at 4
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
CAISI Cyber Assessment
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
reasoningcited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported