GitDealFlowsignals

arXiv preprint · 2022

Constitutional AI: Harmlessness from AI Feedback

Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones

Anthropic

What this paper is

Constitutional AI (Bai et al., 2022) replaces human feedback with a written constitution: the model critiques and revises its own outputs against those principles, then a reward model trained on this AI feedback drives reinforcement learning (RLAIF). Anthropic showed the result is both more helpful and less harmful than RLHF.

Abstract summary

Introduces Constitutional AI (CAI): an alignment approach where an LLM critiques and revises its own outputs according to a written constitution of principles, with reinforcement learning from AI feedback (RLAIF) replacing the human-labeling step. Demonstrates that RLAIF can produce models that are both more helpful AND more harmless than RLHF baselines, while scaling alignment without proportional human labeling effort.

Our summary in our own words, see the canonical source links below for the original abstract.

Why we cite this paper

Constitutional AI is the foundation of Claude's training pipeline at Anthropic, the headline 'safety-first' frontier lab in our engineering-acceleration tracking. The RLAIF paradigm addresses RLHF's scaling bottleneck and has influenced subsequent alignment research across frontier labs.

Where this matters for deal flow

Key findings

  • 1RLAIF (AI feedback) can substitute for RLHF (human feedback) at scale while maintaining alignment quality.
  • 2A written 'constitution' of principles enables transparent control over model behavior.
  • 3Models trained with CAI are both more helpful and more harmless than RLHF baselines on Anthropic's benchmarks.
  • 4Scalable oversight via AI feedback is the path to alignment as models exceed human-evaluator capacity.

Canonical sources

Related glossary terms

Frequently Asked Questions

How does Constitutional AI differ from RLHF?

Standard RLHF uses human preference data to train a reward model; Constitutional AI uses AI-generated preferences against a written constitution. Both pipelines produce aligned models; CAI scales without proportional human-labeling effort.

Who published the Constitutional AI paper?

Anthropic. Authors include Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, and Andy Jones (arXiv:2212.08073, 2022).

What is RLAIF?

Reinforcement Learning from AI Feedback, the technique introduced in this paper, where AI-generated preferences judged against a written constitution replace human preference labeling, allowing alignment to scale without proportional human effort.

Which model uses Constitutional AI?

It is the foundation of Anthropic's Claude training pipeline. The RLAIF approach has also influenced subsequent alignment research across other frontier labs.

See who is building on this, before the round prices it in

We track engineering acceleration across the AI & Machine Learning sector this paper informs: commit velocity, contributor influx, and repo-creation pulse, surfacing breakout teams 21 to 47 days before the fundraise is public.

Five breakout startups, every Sunday, before the round gets crowded

The free Acceleration Watch: five venture-backed teams accelerating on the engineering signal, translated into plain English, 21 to 47 days before the deck circulates. No code-reading, no card.

Signed The Data Nerd · pseudonymous narrator · methodology over personality

Other research papers

Read our own methodology paper

Code-Side Sourcing methodology, replicable on the open dataset.

Read /methodology

Related papers

🚀 Explore Our Network

21-47 days
Signal Lead Time (median 31d)
$80M+
Rounds Tracked
90 sec
Per Scan
5,000+
Founders Tracked

One missed signal is a missed round. Get the Velocity Verdict in your inbox every Sunday free.

Get Free Signals

Free weekly digest. Cancel anytime. No spam, no VC pitches just data.