An alignment technique developed by Anthropic that gives an AI model a set of explicit principles (a "constitution") and has it critique and revise its own outputs against those principles. Complements RLHF by reducing reliance on human labelers for every possible harmful output. the model's behavior is shaped by Constitutional AI alongside RLHF.
Constitutional AI
Also called: CAI
Related terms
This definition is part of a free, structured course on how AI actually works.
Start learning