The Five Core Problems in AI Alignment
A clear-eyed assessment of where alignment research stands today, what remains unsolved, and why we believe the problem is tractable given sufficient focus and talent.
Alignment is not a single problem but a cluster of deeply related challenges. Here we outline our current understanding of the five core problems and our approach to each.
1. Reward Misspecification
The classic problem: optimize for the wrong thing, get the wrong thing. Modern language models are trained on human feedback, but human feedback is noisy, inconsistent, and can be gamed. Constitutional AI represents our approach to this problem.
2. Inner Alignment
Even if the training objective is correctly specified, the model may develop internal representations that pursue different goals. Mechanistic interpretability research is our primary tool for detecting and addressing this.
3. Scalable Oversight
As models become more capable than their supervisors in specific domains, how do we ensure safe behavior? We’re investigating debate, amplification, and recursive reward modeling as potential approaches.
4. Robustness Under Distribution Shift
Models trained in one environment may behave unpredictably in novel contexts. Our safety evaluations specifically probe out-of-distribution behavior.
5. Corrigibility and Shutdown
Will advanced AI systems resist shutdown or modification? We study the conditions under which systems remain corrigible even under capability scaling.