Overview
Superposition is a phenomenon where neural networks represent more features than they have dimensions. This has profound implications for interpretability research, as it means we cannot simply read off features from individual neurons.
Key Insights
- Networks can represent n features in d dimensions where n >> d
- Features exist in superposition when they are sparse and occur rarely
- Superposition creates interference patterns that can be mathematically characterized
- Understanding superposition is a prerequisite for full mechanistic interpretability