We introduce a method for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The core idea is to use a list of rules about harmlessness—a constitution—to auto...
Dr. Sarah Chen — Marcus Webb — Dr. Aisha Patel — James Liu — Dr. Elena Vasquez
As AI systems approach and potentially exceed human-level capabilities in consequential domains, robust safety evaluations become critical. We present a framework for assessing catastrophic risks including CBRN uplift, a...
Dr. Elena Vasquez — Dr. Sarah Chen — Thomas Nakamura — Dr. Marcus Webb
We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training. Some architectural deta...
Neural networks often represent more features than they have dimensions. We call this phenomenon superposition and explore it in toy models, showing how features can be packed into fewer dimensions through interference p...
Dr. Aisha Patel — Dr. James Liu — Dr. Elena Vasquez