Model Cards / Anthropic

Constitutional AI Paper

constitution19,142 words·83 min read·Aug 20, 2026·Source
Version History
Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
19,142-word document condensed to 134 words. Anthropic · Aug 20, 2026
TL;DR

Constitutional AI: Harmlessness from AI Feedback Yuntao Bai∗, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndo

Capability claim
  • We trained for one epoch, using a constant learning rate of 0.5 relative to the pre-training learning rate, and batch size 1024 sequences.
Deployment scope
  • available on BIG Bench [Srivastava et al., 2022].
Limitations the lab flags
  • future work, since training intensively for harmlessness would otherwise result in a model that simply refuses to be helpful.
  • future work. 4In some contexts this could be a virtue [Xu et al., 2020], but in this paper we view it as a problem since it reduces transparency and helpfulness.

Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original constitution · source SHA bb68a3ac3ccc · version dated Aug 20, 2026.