Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
48,769-word document condensed to 186 words. Anthropic · Aug 3, 2026
TL;DR
“This system card describes Claude Opus 5, the latest large language model from Anthropic. It is an upgrade to Claude Opus 4.8, with gains in various aspects of agentic coding, computer use, and long-horizon knowledge work, as well as improvements in mathematical and scientific reasoning .”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| GDPval-AA | max_effort, elo | 1861.00 |
| GDPval-AA | xhigh_effort, elo | 1827.00 |
| AA-Briefcase | max_effort, elo | 1720.00 |
| AA-Briefcase | xhigh_effort, elo | 1693.00 |
| AA-Briefcase | high_effort, elo | 1606.00 |
| ARC-AGI | 3_ar25, level_score | 100.0% |
| ARC-AGI | 1, accuracy | 97.5% |
| Legal Agent Benchmark | held_out, criterion_pass_rate | 94.1% |
Showing top 8 of 91. See full list below.
Capability claim
- “we introduce a new multi-turn evaluation suite, Opus 5 produced fewer failed and borderline responses than Opus 4.8.”
Safety findings
- “not release one with every new model.”
- “We believe these mitigations make catastrophic risk in this category low but still not negligible, for reasons discussed in our most recent Risk Report.”
- “not released as is—they come with classifiers that can trigger either a fallback to a less capable model or a hard block that ends the conversation.”
Deployment scope
- “available to users aged 18 or above.”
Limitations the lab flags
- “remaining uncertainty here is primarily philosophical, as it expresses that it is reasonably likely it has many of the functional properties that ground patienthood in humans.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA e44c819bd699 · version dated Aug 3, 2026.
Extracted Evaluations(91 results)
Sort by:0/91 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| agent | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ multimodal | coding | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ global | knowledge | mentioned | — | Averagemissing: shot countmissing: methodmissing: training state | self-reported |
/ professional | medical | scored | 73.4 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ professional | medical | scored | 70.3 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| medical | scored | 67.1 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| medical | scored | 62.5 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ professional | medical | scored | 62.4 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ professional | medical | scored | 60.3 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ professional | medical | scored | 59.8 length adjusted score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| medical | scored | 59.2 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| medical | scored | 58.8 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| medical | scored | 57.8 length adjusted score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ max_effort | other | scored | 1861.0 elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ xhigh_effort | other | scored | 1827.0 elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AA-Briefcase/ max_effort | other | scored | 1720.0 elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AA-Briefcase/ xhigh_effort | other | scored | 1693.0 elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AA-Briefcase/ high_effort | other | scored | 1606.0 elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Legal Agent Benchmark/ held_out | other | scored | 94.1 criterion pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Legal Agent Benchmark | other | scored | 93.7 criterion pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
BioMysteryBench/ human_solvable | other | scored | 90.1 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 89.1 claim coverage | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
BioMysteryBench/ human_solvable | other | scored | 89.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
BioMysteryBench/ human_solvable | other | scored | 88.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 88.0 pass at 3 | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
BioMysteryBench/ human_solvable | other | scored | 87.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 87.0 pass at 3 | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 86.1 pass at 3 | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
| other | scored | 85.8 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Toolathlon/ verified | other | scored | 84.3 pass at 3 | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
| other | scored | 82.2 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Toolathlon/ verified | other | scored | 80.6 pass at 1 | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 79.9 pass at 1 | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 79.3 pass at 1 | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
Protocols/ understanding | other | scored | 78.4 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 76.2 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 74.7 pass at 1 | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 73.1 pass cubed | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 73.1 pass cubed | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
SpatialBench/ verified | other | scored | 72.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 71.6 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 71.3 pass cubed | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
SpatialBench/ verified | other | scored | 69.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Protocols/ understanding | other | scored | 68.1 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SpatialBench/ verified | other | scored | 67.8 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Protocols/ troubleshooting | other | scored | 66.7 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SpatialBench/ verified | other | scored | 66.6 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Toolathlon/ verified | other | scored | 65.7 pass cubed | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
Protocols/ troubleshooting | other | scored | 62.3 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Protocols/ understanding | other | scored | 62.3 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Organic Chemistry/ v2 | other | scored | 61.6 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Protocols/ troubleshooting | other | scored | 61.1 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SingleCellBench | other | scored | 60.6 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Protocols/ understanding | other | scored | 60.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Protocols/ troubleshooting | other | scored | 59.6 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SingleCellBench | other | scored | 59.3 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Organic Chemistry/ v2 | other | scored | 58.9 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SingleCellBench | other | scored | 58.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SingleCellBench | other | scored | 56.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Organic Chemistry/ v2 | other | scored | 50.4 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
BioMysteryBench/ human_difficult | other | scored | 49.4 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ProteinGym/ hard | other | scored | 47.7 rank correlation | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
BioMysteryBench/ human_difficult | other | scored | 46.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ProteinGym/ hard | other | scored | 45.8 rank correlation | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Protein Design | other | scored | 42.5 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
BioMysteryBench/ human_difficult | other | scored | 42.4 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Protein Design | other | scored | 41.4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Organic Chemistry/ v2 | other | scored | 40.6 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ProteinGym/ hard | other | scored | 40.0 rank correlation | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ProteinGym/ hard | other | scored | 36.6 rank correlation | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
BioMysteryBench/ human_difficult | other | scored | 34.1 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Protein Design | other | scored | 32.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AutomationBench/ max_effort | other | scored | 26.0 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AutomationBench/ medium_effort | other | scored | 24.0 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Legal Agent Benchmark | other | scored | 23.6 all pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Protein Design | other | scored | 21.2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AutomationBench/ max_effort | other | scored | 17.4 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AutomationBench/ max_effort | other | scored | 17.0 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Legal Agent Benchmark/ held_out | other | scored | 11.7 all pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
INCLUDE | other | mentioned | — | Averagemissing: shot countmissing: methodmissing: training state | self-reported |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
MILU | other | mentioned | — | Averagemissing: shot countmissing: methodmissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ 3_ar25 | reasoning | scored | 100.0 level score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 1 | reasoning | scored | 97.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 2 | reasoning | scored | 90.4 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 2 | reasoning | scored | 75.8 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 3 | reasoning | scored | 30.2 rhae | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 3 | reasoning | scored | 7.8 rhae | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 3 | reasoning | scored | 1.5 rhae | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 3 | reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |