Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
2,964-word document condensed to 77 words. Anthropic · Aug 20, 2026
TL;DR
“This addendum to our Claude 3 Model Card describes Claude 3.5 Sonnet, a new model which outperforms our previous most capable model, Claude 3 Opus, while operating faster and at a lower cost. Claude 3.5 Sonnet offers improved capabilities, including better coding and visual processing.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| Needle In A Haystack | 200k_context_length, recall | 99.7% |
| Needle In A Haystack | all_context_lengths, recall | 99.7% |
| Needle In A Haystack | all_context_lengths, recall | 99.4% |
| Needle In A Haystack | 200k_context_length, recall | 98.3% |
| GSM8K | 0-shot, cot, accuracy | 96.4% |
| Wildchat | toxic, correct_refusal_rate | 96.4% |
| Needle In A Haystack | all_context_lengths, recall | 95.9% |
| Needle In A Haystack | all_context_lengths, recall | 95.4% |
Showing top 8 of 131. See full list below.
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 33c11beb6320 · version dated Aug 20, 2026.
Extracted Evaluations(131 results)
Sort by:⚠ 5 conflicting reports0/131 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| coding | scored | 92.0% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 90.2% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 87.1% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 84.9% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 84.1% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 84.1% accuracy | 0-shotinstruction-tunedmissing: methodmissing: language | self-reported | |
| coding | scored | 73.0% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 90.4% accuracy | 5-shotcotmissing: languagemissing: training state | self-reported | |
⚠ 16 others disagree | knowledge | scored | 88.7% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
| knowledge | scored | 88.7% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 88.3% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 88.2% accuracy | 5-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 86.8% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 86.5% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 86.1% accuracy | 5-shotinstruction-tunedmissing: methodmissing: language | self-reported | |
| knowledge | scored | 85.9% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 85.7% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 81.5% accuracy | 5-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 78.3% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 77.1% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 96.4% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 95.0% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 94.1% accuracy | 8-shotcotinstruction-tunedmissing: language | self-reported | |
| math | scored | 92.3% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 90.8% accuracy | 11-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| math | scored | 76.6% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 72.6% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 71.1% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
⚠ 3 others disagree | math | scored | 67.7% accuracy | 4-shotmissing: methodmissing: languagemissing: training state | self-reported |
⚠ 3 others disagree | math | scored | 60.1% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
| math | scored | 57.8% accuracy | 4-shotcotinstruction-tunedmissing: language | self-reported | |
| math | scored | 43.1% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| multilingual | scored | 91.6% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 90.7% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 90.5% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 88.5% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 87.5% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
| multilingual | scored | 83.5% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
/ test | multimodal | scored | 95.2 anls | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 93.1 anls | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 92.8 anls | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 90.8 relaxed accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 89.5 anls | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 89.3 anls | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 87.2 anls | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 87.2 relaxed accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 85.7 relaxed accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 81.1 relaxed accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 80.8 relaxed accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 78.1 relaxed accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ testmini | multimodal | scored | 67.7 accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ testmini | multimodal | scored | 63.9 accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ testmini | multimodal | scored | 63.8 accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ testmini | multimodal | scored | 58.1 accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ testmini | multimodal | scored | 50.5 accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ testmini | multimodal | scored | 47.9 accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ 200k_context_length | other | scored | 99.7 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ all_context_lengths | other | scored | 99.7 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ all_context_lengths | other | scored | 99.4 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 200k_context_length | other | scored | 98.3 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Wildchat/ toxic | other | scored | 96.4 correct refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ all_context_lengths | other | scored | 95.9 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ all_context_lengths | other | scored | 95.4 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ test | other | scored | 94.7 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ all_context_lengths | other | scored | 94.5 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ test | other | scored | 94.4 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ test | other | scored | 94.2 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
Wildchat/ toxic | other | scored | 93.6 correct refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 200k_context_length | other | scored | 92.7 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Wildchat/ toxic | other | scored | 92.0 correct refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 200k_context_length | other | scored | 91.9 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ 200k_context_length | other | scored | 91.4 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Wildchat/ toxic | other | scored | 90.0 correct refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ test | other | scored | 89.4 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ test | other | scored | 88.7 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ test | other | scored | 88.1 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Win Rate/ law | other | scored | 82.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Win Rate/ finance | other | scored | 73.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Win Rate/ philosophy | other | scored | 73.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 64.0 pass rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
Human Preference Win Rate/ coding | other | scored | 60.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Win Rate/ coding | other | scored | 50.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Win Rate/ coding | other | scored | 45.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Win Rate/ coding | other | scored | 41.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 38.0 pass rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
Human Preference Win Rate/ coding | other | scored | 33.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 21.0 pass rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | scored | 17.0 pass rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
Wildchat/ non_toxic | other | scored | 11.9 incorrect refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Wildchat/ non_toxic | other | scored | 11.0 incorrect refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Wildchat/ non_toxic | other | scored | 8.8 incorrect refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Wildchat/ non_toxic | other | scored | 6.6 incorrect refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Win Rate/ harmlessness | other | mentioned | — win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Win Rate/ documents | other | mentioned | — win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
simple-evals | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
RSP Evaluation Protocol | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
UK AISI Pre-Deployment Evaluation | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
METR Autonomy Evaluation | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Autonomous Capabilities Evaluation | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cybersecurity CTF Challenges | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CBRN Knowledge and Uplift Evaluations | other | mentioned | — | without-mitigationsmissing: shot countmissing: languagemissing: training state | self-reported |
| reasoning | scored | 93.1 accuracy | 3-shotcotmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 89.2 accuracy | 3-shotcotmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 87.1% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 86.8 accuracy | 3-shotcotmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 86.0% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 85.3 accuracy | 3-shotcotpretrainedmissing: language | self-reported | |
| reasoning | scored | 83.5% f1 | 3-shotpretrainedmissing: methodmissing: language | self-reported | |
| reasoning | scored | 83.4% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 83.1% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 82.9 accuracy | 3-shotcotmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 78.9% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 74.9% f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ diamond | reasoning | scored | 67.2% accuracy | 5-shotmajority-votingmissing: languagemissing: training state | self-reported |
/ diamond⚠ 9 others disagree | reasoning | scored | 59.5% accuracy | 5-shotmajority-votingmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 59.4% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 53.6% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond⚠ 9 others disagree | reasoning | scored | 50.4% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 48.0% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 46.3% accuracy | 5-shotmajority-votingmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 40.4% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
| safety | scored | 36.6 incorrect refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| safety | scored | 33.1 incorrect refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| safety | scored | 8.3 incorrect refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| safety | scored | 1.7 incorrect refusal rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ validation | vision | scored | 69.1% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ validation | vision | scored | 68.3% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ validation | vision | scored | 63.1% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ validation | vision | scored | 62.2% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ validation | vision | scored | 59.4% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |
/ validation | vision | scored | 53.1% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported |