AXIOM
Get API Access
All Research
ScalingLanguage ModelsComputeResearch

Scaling Laws for Neural Language Models

We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training. Some architectural details such as network width or depth have minimal effects within a wide range. These results allow us to determine the optimal allocation of a fixed compute budget.

Published
Authors
Dr. Marcus WebbDr. Sarah ChenRyan Park

Abstract

We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude.

Key Findings

  • Model performance improves predictably with scale
  • Optimal allocation of compute budget between parameters and training tokens
  • Architecture details have minimal effect on scaling behavior
  • Data quality matters more than quantity above certain thresholds