ScalingLanguage ModelsComputeResearch
Scaling Laws for Neural Language Models
We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training. Some architectural details such as network width or depth have minimal effects within a wide range. These results allow us to determine the optimal allocation of a fixed compute budget.
Abstract
We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude.
Key Findings
- Model performance improves predictably with scale
- Optimal allocation of compute budget between parameters and training tokens
- Architecture details have minimal effect on scaling behavior
- Data quality matters more than quantity above certain thresholds