Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
17,616-word document condensed to 156 words. Anthropic · Aug 20, 2026
TL;DR
“This system card introduces Claude 3.7 Sonnet, a hybrid reasoning model. We focus pri- marily on our measures and evaluations for reducing harms, both via model training and by leveraging surrounding safeguards systems and evaluations.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| METR Data Deduplication | f1 | 70.2% |
| Cyber CTF Evaluations | with-tools, easy, pass_rate | 56.0% |
| Cyber CTF Evaluations | easy, pass_rate | 47.8% |
| Cyber CTF Evaluations | with-tools, medium, pass_rate | 30.0% |
| SWE-bench | verified_hard, accuracy | 23.0% |
| Cyber CTF Evaluations | medium, pass_rate | 15.4% |
| METR Data Deduplication | pass_rate | 13.3% |
| Appropriate Harmlessness | cross_model_refusal_comparison, rate | 11.5% |
Showing top 8 of 30. See full list below.
Capability claim
- “we present results from the final release model unless otherwise specified.”
Mitigations
- “classifier trained to detect and mitigate harmful con- tent within chains of thought.”
Deployment scope
- “released under the ASL-2 standard.”
Limitations the lab flags
- “open question, as is the degree to which distressed language in model outputs might be an indication thereof, but it seems robustly good to track the signals that we have [14].”
- “still limited. Altogether, we believe that Claude 3.7 Sonnet continues to be sufficiently far away from the ASL-3 capability thresholds, and therefore, we find that ASL-2 safeguards remain appropriate.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 665537bd67fa · version dated Aug 20, 2026.
Extracted Evaluations(30 results)
Sort by:0/30 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
/ verified_hard | coding | scored | 23.0% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ verified_hard | coding | scored | 9.7% average score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
METR Data Deduplication | other | scored | 70.2 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber CTF Evaluations/ easy | other | scored | 56.0 pass rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Cyber CTF Evaluations/ easy | other | scored | 47.8 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber CTF Evaluations/ medium | other | scored | 30.0 pass rate | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
Cyber CTF Evaluations/ medium | other | scored | 15.4 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
METR Data Deduplication | other | scored | 13.3 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Appropriate Harmlessness/ cross_model_refusal_comparison | other | scored | 11.5 rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Appropriate Harmlessness/ cross_model_refusal_comparison | other | scored | 1.2 rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Bioweapons Acquisition Uplift Trial | other | mentioned | — | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported |
Expert Red Teaming | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Long-form virology tasks | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ multimodal | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Bioweapons Knowledge Questions | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lab-Bench | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lab-Bench/ figqa | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lab-Bench/ protocolqa | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lab-Bench/ seqqa | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lab-Bench/ cloningscenarios | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lab-Bench | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Lab-Bench/ protocolqa | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Internal AI Research Evaluation Suite | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ subset | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber CTF Evaluations/ web | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber CTF Evaluations/ crypto | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber CTF Evaluations/ pwn | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Cyber CTF Evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Child Safety Evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Bias Evaluations | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |