AXIOM Research Publications Volume I · 4 Papers · 2024—2026

Research

Alignment · Compute · Constitutional AI · Evaluation · Features · Frontier AI · Interpretability · Language Models · Mechanistic Analysis · Neural Networks · RLHF · Red-Teaming · Research · Safety · Scaling

Index

Safety · RLHF · Constitutional AI · Alignment Featured

Constitutional AI: Harmlessness from AI Feedback

We introduce a method for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The core idea is to use a list of rules about harmlessness—a constitution—to auto…

Safety · Evaluation · Red-Teaming · Frontier AI

Frontier Safety Evaluations: A Framework for Catastrophic Risk Assessment

As AI systems approach and potentially exceed human-level capabilities in consequential domains, robust safety evaluations become critical. We present a framework for assessing catastrophic risks including CBRN uplift, a…

Scaling · Language Models · Compute · Research

Scaling Laws for Neural Language Models

We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training. Some architectural deta…

Interpretability · Mechanistic Analysis · Neural Networks · Features Featured

Toy Models of Superposition

Neural networks often represent more features than they have dimensions. We call this phenomenon superposition and explore it in toy models, showing how features can be packed into fewer dimensions through interference p…