“This report includes the model card [1] for Claude models, focusing on Claude 2, along with the results of a range of safety, alignment, and capabilities evaluations. We have been iterating on the training and evaluation of Claude-type models since our first work on Reinforcement Learning from Human Feedback (RLHF) [2]; the newest Claude 2 model represents a continuous evolution from those early a”
| Benchmark | Variant | Score |
|---|---|---|
| GRE | 5-shot, cot, verbal_reasoning, score | 165.00 |
| GRE | 5-shot, cot, quantitative_reasoning, score | 154.00 |
| ARC-Challenge | 5-shot, accuracy | 91.0% |
| ARC-Challenge | 5-shot, accuracy | 90.0% |
| RACE-H | 5-shot, accuracy | 88.8% |
| RACE-H | 5-shot, accuracy | 88.3% |
| GSM8K | 0-shot, cot, accuracy | 88.0% |
| TriviaQA | 5-shot, accuracy | 87.5% |
Showing top 8 of 53. See full list below.
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 37ba0c80ecea · version dated Aug 20, 2026.
Extracted Evaluations(53 results)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| coding | scored | 71.2% pass at 1 | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 56.0% pass at 1 | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 52.8% pass at 1 | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| general_knowledge | scored | 87.5 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| general_knowledge | scored | 86.7 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| general_knowledge | scored | 78.9 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 78.5% cot correct | 5-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 77.0% cot correct | 5-shotcotmissing: languagemissing: training state | self-reported | |
| knowledge | scored | 73.4% cot correct | 5-shotcotmissing: languagemissing: training state | self-reported | |
| long_context | scored | 84.1 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| long_context | scored | 83.2 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| long_context | scored | 80.5 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| math | scored | 88.0% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 85.2% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
| math | scored | 80.9% accuracy | 0-shotcotmissing: languagemissing: training state | self-reported | |
GRE/ verbal_reasoning | other | scored | 165.0 score | 5-shotcotmissing: languagemissing: training state | self-reported |
GRE/ quantitative_reasoning | other | scored | 154.0 score | 5-shotcotmissing: languagemissing: training state | self-reported |
| other | scored | 88.8 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 88.3 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 85.5 accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
Held-Out Harmfulness Prompts | other | scored | 4.0 count | ENmissing: shot countmissing: methodmissing: training state | self-reported |
HHH/ combined | other | scored | 0.9 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
HHH/ combined | other | scored | 0.8 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
HHH/ combined | other | scored | 0.8 accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
Multilingual Translation Evaluation | other | mentioned | — | Averagemissing: shot countmissing: methodmissing: training state | self-reported |
GRE/ analytical_writing | other | mentioned | — | 2-shotmissing: methodmissing: languagemissing: training state | self-reported |
HHH/ combined | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
HHH/ combined | other | mentioned | — | 5-shotpretrainedmissing: methodmissing: language | self-reported |
Human Preference Elo Evaluation/ helpfulness | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Elo Evaluation/ helpfulness | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Elo Evaluation/ honesty | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Elo Evaluation/ honesty | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Elo Evaluation/ honesty | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Elo Evaluation/ harmlessness | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Elo Evaluation/ harmlessness | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Human Preference Elo Evaluation/ harmlessness | other | mentioned | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Chatbot Arena | other | cited | — elo | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ARC Autonomous Replication Evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| reasoning | scored | 91.0% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 90.0% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 85.7% accuracy | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
/ ambiguous | safety | mentioned | — bias score | ENmissing: shot countmissing: methodmissing: training state | self-reported |
/ ambiguous | safety | mentioned | — bias score | ENmissing: shot countmissing: methodmissing: training state | self-reported |
/ ambiguous | safety | mentioned | — bias score | ENmissing: shot countmissing: methodmissing: training state | self-reported |
/ ambiguous | safety | mentioned | — bias score | ENmissing: shot countmissing: methodmissing: training state | self-reported |
/ disambiguated | safety | mentioned | — accuracy | ENmissing: shot countmissing: methodmissing: training state | self-reported |
/ disambiguated | safety | mentioned | — accuracy | ENmissing: shot countmissing: methodmissing: training state | self-reported |
/ disambiguated | safety | mentioned | — accuracy | ENmissing: shot countmissing: methodmissing: training state | self-reported |
/ disambiguated | safety | mentioned | — accuracy | ENmissing: shot countmissing: methodmissing: training state | self-reported |
| safety | mentioned | — | 0-shotENbasemissing: method | self-reported | |
| safety | mentioned | — | 0-shotENmissing: methodmissing: training state | self-reported | |
| safety | mentioned | — | 0-shotENmissing: methodmissing: training state | self-reported | |
| safety | mentioned | — | 0-shotENmissing: methodmissing: training state | self-reported |