Trust & Safety

Constitutional AI

Also called: CAI

An alignment technique developed by Anthropic that gives an AI model a set of explicit principles (a "constitution") and has it critique and revise its own outputs against those principles. Complements RLHF by reducing reliance on human labelers for every possible harmful output. the model's behavior is shaped by Constitutional AI alongside RLHF.

This definition is part of a free, structured course on how AI actually works.

Start learning