“Constitutional AI: Harmlessness from AI Feedback Yuntao Bai∗, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndo”
- “We trained for one epoch, using a constant learning rate of 0.5 relative to the pre-training learning rate, and batch size 1024 sequences.”
- “available on BIG Bench [Srivastava et al., 2022].”
- “future work, since training intensively for harmlessness would otherwise result in a model that simply refuses to be helpful.”
- “future work. 4In some contexts this could be a virtue [Xu et al., 2020], but in this paper we view it as a problem since it reduces transparency and helpfulness.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original constitution · source SHA bb68a3ac3ccc · version dated Aug 20, 2026.