Abstract
We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude.
Key Findings
- Model performance improves predictably with scale
- Optimal allocation of compute budget between parameters and training tokens
- Architecture details have minimal effect on scaling behavior
- Data quality matters more than quantity above certain thresholds