Scaling laws
The empirical finding that model loss falls predictably with more compute, parameters, and data — and (via Chinchilla) that smaller models trained on more tokens often beat bigger undertrained ones.
The empirical finding that model loss falls predictably with more compute, parameters, and data — and (via Chinchilla) that smaller models trained on more tokens often beat bigger undertrained ones.