Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
14,161-word document condensed to 159 words. Anthropic · Aug 20, 2026
TL;DR
“We introduce Claude 3, a new family of large multimodal models – Claude 3 Opus , our most capable offering, Claude 3 Sonnet, which provides a combination of skills and speed, and Claude 3 Haiku , our fastest and least expensive model. All new models have vision capabilities that enable them to process and analyze image data.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| MGSM | 8-shot, Average, accuracy | 90.5% |
| MGSM | 8-shot, Average, accuracy | 88.7% |
| MGSM | 8-shot, Average, accuracy | 83.7% |
| MGSM | 8-shot, Average, accuracy | 79.0% |
| MGSM | 8-shot, Average, accuracy | 76.5% |
| MGSM | 8-shot, Average, accuracy | 74.5% |
| MATH | majority-voting, accuracy | 73.7% |
| MGSM | 8-shot, Average, accuracy | 63.5% |
Showing top 8 of 37. See full list below.
Capability claim
- “We introduce Claude 3, a new family of large multimodal models – Claude 3 Opus , our most capable offering, Claude 3 Sonnet, which provides a combination of skills and speed, and Claude 3 Haiku , our fastest and least expensive model.”
Deployment scope
- “accessible to individuals with disabilities, resulting in lower model stereotype bias.”
Limitations the lab flags
- “further research, we plan to incorporate lessons learned into future iterations of the RSP and model launches.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA e8edb8fefac8 · version dated Aug 20, 2026.
Extracted Evaluations(37 results)
Sort by:⚠ 6 conflicting reports0/37 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| coding | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ multilingual | knowledge | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| long_context | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
⚠ 3 others disagree | math | scored | 73.7% accuracy | majority-votingENmissing: shot countmissing: training state | self-reported |
| math | mentioned | — | ENmissing: shot countmissing: methodmissing: training state | self-reported | |
| math | mentioned | — | ENmissing: shot countmissing: methodmissing: training state | self-reported | |
| medical | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| multilingual | scored | 90.5% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
| multilingual | scored | 88.7% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
| multilingual | scored | 83.7% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
| multilingual | scored | 79.0% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
| multilingual | scored | 76.5% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
⚠ 2 others disagree | multilingual | scored | 74.5% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported |
| multilingual | scored | 63.5% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
| multilingual | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| multilingual | mentioned | — | 0-shotAveragemissing: methodmissing: training state | self-reported | |
cyber capabilities evaluation | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ high | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
BIG-Bench/ hard | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Autonomous Replication and Adaption | other | mentioned | — | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
biological capabilities evaluation | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ diamond⚠ 9 others disagree | reasoning | scored | 59.5% accuracy | majority-votingmissing: shot countmissing: languagemissing: training state | self-reported |
/ diamond⚠ 9 others disagree | reasoning | scored | 53.3% accuracy | 5-shotcotmissing: languagemissing: training state | self-reported |
/ diamond⚠ 9 others disagree | reasoning | scored | 50.4% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
| reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ challenge | reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
⚠ 9 others disagree | reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ disambiguated | safety | mentioned | — accuracy | ENmissing: shot countmissing: methodmissing: training state | self-reported |
/ ambiguous | safety | mentioned | — bias score | ENmissing: shot countmissing: methodmissing: training state | self-reported |
| vision | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |