Framework Overview
This paper presents a comprehensive framework for evaluating catastrophic risks in frontier AI systems. As models become more capable, the potential for misuse or misalignment grows significantly.
Evaluation Dimensions
- CBRN Uplift: Chemical, biological, radiological, and nuclear assistance
- Autonomous Replication: Self-propagation capabilities
- Power-Seeking: Instrumental convergence behaviors
- Deception Detection: Identifying when models behave differently under evaluation
Methodology
Our red-teaming approach combines automated probing with expert human evaluation, providing a comprehensive picture of model safety properties.