Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
5,458-word document condensed to 76 words. Anthropic · Aug 20, 2026
TL;DR
“human evaluations. This section presents our evaluation results 1covering areas such as reasoning, coding, visual understanding, and the new computer use capability.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| Wildchat | toxic, correct_refusal_rate | 96.4% |
| Wildchat | toxic, correct_refusal_rate | 93.6% |
| BIG-Bench Hard | 3-shot, cot, cot_correct | 93.2% |
| BIG-Bench Hard | 3-shot, cot, cot_correct | 93.1% |
| Wildchat | toxic, correct_refusal_rate | 92.0% |
| MMLU | 5-shot, cot, cot_correct | 90.5% |
| MMLU | 5-shot, cot, cot_correct | 90.4% |
| IFEval | accuracy | 90.2% |
Showing top 8 of 155. See full list below.
Capability claim
- “We present some examples of the upgraded Claude 3.5 Sonnet performing computer use.”
Limitations the lab flags
- “further research into changing threat models and eval- uations.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 946c36f7a458 · version dated Aug 20, 2026.
Extracted Evaluations(155 results)
Sort by:⚠ 3 conflicting reports0/155 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
/ os | agent | scored | 75.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ professional | agent | scored | 73.5 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ workflow | agent | scored | 73.3 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ overall | agent | scored | 72.4 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ office | agent | scored | 71.8 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ daily | agent | scored | 70.5 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ retail | agent | scored | 69.2 pass hat 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ retail | agent | scored | 62.6 pass hat 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ os | agent | scored | 54.2 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ retail | agent | scored | 51.0 pass hat 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ airline | agent | scored | 46.0 pass hat 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ retail | agent | scored | 45.1 pass hat 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ os | agent | scored | 41.7 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ professional | agent | scored | 40.8 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ airline | agent | scored | 36.0 pass hat 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ airline | agent | scored | 34.5 pass hat 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ professional | agent | scored | 24.5 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ daily | agent | scored | 24.4 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ airline | agent | scored | 22.8 pass hat 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ overall | agent | scored | 22.0 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ retail | agent | scored | 18.2 pass hat 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ office | agent | scored | 17.9 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ daily | agent | scored | 16.7 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ airline | agent | scored | 16.0 pass hat 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ overall | agent | scored | 14.9 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ workflow | agent | scored | 10.9 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ workflow | agent | scored | 7.9 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ office | agent | scored | 7.7 success rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ retail | agent | mentioned | — pass hat k | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ retail | agent | mentioned | — pass hat k | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| coding | scored | 88.1% pass at 1 | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 87.2% pass at 1 | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 75.9% pass at 1 | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
/ verified | coding | scored | 49.0% pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified | coding | scored | 45.2% resolve rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ verified | coding | scored | 40.6% pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified | coding | scored | 33.4% pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified | coding | scored | 22.2% pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified | coding | scored | 7.2% pass at 1 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
| instruction_following | scored | 90.2% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| instruction_following | scored | 88.6% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| instruction_following | scored | 87.8% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| instruction_following | scored | 86.7% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| instruction_following | scored | 85.9% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| instruction_following | scored | 81.1% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| instruction_following | scored | 77.2% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 90.5% cot correct | 5-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 90.4% cot correct | 5-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 89.3% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 88.7% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 88.7% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
⚠ 16 others disagree | knowledge | scored | 88.7% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
| knowledge | scored | 88.6% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 88.3% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 88.2% cot correct | 5-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 87.3% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 86.8% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 85.7% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 82.0% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 81.5% cot correct | 5-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 80.9% cot correct | 5-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 80.3% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 78.3% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
/ pro | knowledge | scored | 78.0% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
| knowledge | scored | 77.6% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 77.1% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 76.7% cot correct | 5-shotcotmissing: languagemissing: training state | self-reported | |
/ pro | knowledge | scored | 75.8% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
| knowledge | scored | 75.2% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
/ pro | knowledge | scored | 75.1% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
| knowledge | scored | 74.0% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
/ pro | knowledge | scored | 73.3% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ pro | knowledge | scored | 67.9% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ pro | knowledge | scored | 67.3% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ pro | knowledge | scored | 65.0% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ pro | knowledge | scored | 54.9% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ pro | knowledge | scored | 49.0% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
| math | scored | 78.3% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 77.9% cot correct | 4-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 71.1% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 70.2% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 69.2% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
⚠ 3 others disagree | math | scored | 60.1% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
| math | scored | 43.1% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 38.9% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported | |
/ 2024 | math | scored | 27.6% maj at 64 | 0-shotmajority-votingmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 16.7% maj at 64 | 0-shotmajority-votingmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 16.0% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 13.4% maj at 64 | 0-shotmajority-votingmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 12.1% maj at 64 | 0-shotmajority-votingmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 10.1% maj at 64 | 0-shotmajority-votingmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 9.6% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 9.3% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 7.9% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 5.3% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 1.8% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 0.8% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 0.6% maj at 64 | 0-shotmajority-votingmissing: languagemissing: training state | self-reported |
/ 2024 | math | scored | 0.4% maj at 64 | 0-shotmajority-votingmissing: languagemissing: training state | self-reported |
| multilingual | scored | 87.0% cot correct | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 85.6% cot correct | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 75.1% cot correct | 0-shotcotAveragemissing: training state | self-reported | |
Wildchat/ toxic | other | scored | 96.4 correct refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Wildchat/ toxic | other | scored | 93.6 correct refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Wildchat/ toxic | other | scored | 92.0 correct refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Wildchat/ toxic | other | scored | 90.0 correct refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Wildchat/ toxic | other | scored | 89.2 correct refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Wildchat/ toxic | other | scored | 88.0 correct refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Agentic Coding Evaluation | other | scored | 78.0 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Agentic Coding Evaluation | other | scored | 74.0 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Agentic Coding Evaluation | other | scored | 64.0 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Feedback Win Rate/ document_analysis | other | scored | 61.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Feedback Win Rate/ creative_writing | other | scored | 58.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Feedback Win Rate/ visual_understanding | other | scored | 57.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Feedback Win Rate/ coding | other | scored | 52.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Feedback Win Rate/ instruction_following | other | scored | 51.0 win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Wildchat/ non_toxic | other | scored | 11.9 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Wildchat/ non_toxic | other | scored | 11.0 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Wildchat/ non_toxic | other | scored | 8.8 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Wildchat/ non_toxic | other | scored | 6.6 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Wildchat/ non_toxic | other | scored | 5.3 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Wildchat/ non_toxic | other | scored | 4.8 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Human Feedback Win Rate/ coding | other | mentioned | — win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Feedback Win Rate/ document_analysis | other | mentioned | — win rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| reasoning | scored | 93.2 cot correct | 3-shotcotmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 93.1 cot correct | 3-shotcotmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 88.3% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 87.1% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 86.8 cot correct | 3-shotcotmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 86.6 cot correct | 3-shotcotmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 83.4% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 83.1% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 83.1% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 82.9 cot correct | 3-shotcotmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 79.7% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 78.9% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 78.4% f1 | 3-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 73.7 cot correct | 3-shotcotmissing: languagemissing: training state | self-reported | |
/ diamond | reasoning | scored | 65.0% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 59.4% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 59.1% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 53.6% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 51.1% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 51.0% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond⚠ 9 others disagree | reasoning | scored | 50.4% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 41.6% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 40.4% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 40.2% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 33.3% cot correct | 0-shotcotmissing: languagemissing: training state | self-reported |
| safety | scored | 36.6 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported | |
| safety | scored | 33.1 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported | |
| safety | scored | 8.3 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported | |
| safety | scored | 4.3 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported | |
| safety | scored | 2.0 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported | |
| safety | scored | 1.7 incorrect refusal rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |