AXIOM
Get API Access
All Research
InterpretabilityMechanistic AnalysisNeural NetworksFeatures

Toy Models of Superposition

Neural networks often represent more features than they have dimensions. We call this phenomenon superposition and explore it in toy models, showing how features can be packed into fewer dimensions through interference patterns. Understanding superposition is crucial for mechanistic interpretability.

Published
Authors
Dr. Aisha PatelDr. James LiuDr. Elena Vasquez

Overview

Superposition is a phenomenon where neural networks represent more features than they have dimensions. This has profound implications for interpretability research, as it means we cannot simply read off features from individual neurons.

Key Insights

  1. Networks can represent n features in d dimensions where n >> d
  2. Features exist in superposition when they are sparse and occur rarely
  3. Superposition creates interference patterns that can be mathematically characterized
  4. Understanding superposition is a prerequisite for full mechanistic interpretability