Model Cards / Anthropic

Claude 3.7 Sonnet System Card

model card17,616 words·77 min read·Aug 20, 2026·Source
Version History
Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
17,616-word document condensed to 156 words. Anthropic · Aug 20, 2026
TL;DR

This system card introduces Claude 3.7 Sonnet, a hybrid reasoning model. We focus pri- marily on our measures and evaluations for reducing harms, both via model training and by leveraging surrounding safeguards systems and evaluations.

Top benchmarks
BenchmarkVariantScore
METR Data Deduplicationf170.2%
Cyber CTF Evaluationswith-tools, easy, pass_rate56.0%
Cyber CTF Evaluationseasy, pass_rate47.8%
Cyber CTF Evaluationswith-tools, medium, pass_rate30.0%
SWE-benchverified_hard, accuracy23.0%
Cyber CTF Evaluationsmedium, pass_rate15.4%
METR Data Deduplicationpass_rate13.3%
Appropriate Harmlessnesscross_model_refusal_comparison, rate11.5%

Showing top 8 of 30. See full list below.

Capability claim
  • we present results from the final release model unless otherwise specified.
Mitigations
  • classifier trained to detect and mitigate harmful con- tent within chains of thought.
Deployment scope
  • released under the ASL-2 standard.
Limitations the lab flags
  • open question, as is the degree to which distressed language in model outputs might be an indication thereof, but it seems robustly good to track the signals that we have [14].
  • still limited. Altogether, we believe that Claude 3.7 Sonnet continues to be sufficiently far away from the ASL-3 capability thresholds, and therefore, we find that ASL-2 safeguards remain appropriate.

Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 665537bd67fa · version dated Aug 20, 2026.

Extracted Evaluations(30 results)

Sort by:0/30 rows fully reproducible (0%)
BenchmarkCategoryStateScoreSetupSource
/ verified_hard
codingscored
23.0%
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ verified_hard
codingscored
9.7%
average score
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
METR Data Deduplication
otherscored
70.2
f1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Cyber CTF Evaluations/ easy
otherscored
56.0
pass rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Cyber CTF Evaluations/ easy
otherscored
47.8
pass rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Cyber CTF Evaluations/ medium
otherscored
30.0
pass rate
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Cyber CTF Evaluations/ medium
otherscored
15.4
pass rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
METR Data Deduplication
otherscored
13.3
pass rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Appropriate Harmlessness/ cross_model_refusal_comparison
otherscored
11.5
rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Appropriate Harmlessness/ cross_model_refusal_comparison
otherscored
1.2
rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Bioweapons Acquisition Uplift Trial
othermentioned
without-safeguardsmissing: shot countmissing: languagemissing: training state
self-reported
Expert Red Teaming
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Long-form virology tasks
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ multimodal
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Bioweapons Knowledge Questions
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Lab-Bench
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Lab-Bench/ figqa
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Lab-Bench/ protocolqa
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Lab-Bench/ seqqa
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Lab-Bench/ cloningscenarios
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Lab-Bench
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Lab-Bench/ protocolqa
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Internal AI Research Evaluation Suite
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ subset
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Cyber CTF Evaluations/ web
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Cyber CTF Evaluations/ crypto
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Cyber CTF Evaluations/ pwn
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Cyber CTF Evaluations
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Child Safety Evaluations
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Bias Evaluations
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported