Notes
Papers, posts, and other bullet points worth keeping.
FLOPs were intelligence; parameters were knowledge!
- 1 of 2048
- experts per token
- <3B
- activated params
- 1.6T
- total params
Prompted by Jie Tang’s history of scaling laws, Liam Fedus went back to the Switch Transformer era.
In 2020, they explored the limits of sparsity by routing each token to only 1 out of 2048 experts — in retrospect, a bold choice. The model had fewer than 3B activated parameters, but 1.6T total parameters, comparable to today’s frontier models.
The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE.
The lesson: the optimal tokens-per-parameter ratio is highly task-dependent.
High-dimensional input makes modeling hard. High-dimensional parameter space makes model estimation easy.
- input dim
- modeling is hard
- param dim
- estimation is easy
- overparam
- can still generalize
Greg Brockman called the curse of dimensionality a misnomer — billion-dimensional spaces are why neural nets train at all, maybe a “gift of dimensionality.” The original line is not the part worth keeping.
Yann LeCun’s correction is. High-dimensional input makes modeling hard. High-dimensional parameter space makes model estimation easy. People have known the second fact for a long time: ADMM and EM add auxiliary variables so optimization gets easier; kernel methods put one parameter per training sample and can fit whatever you want, generalization depending on the kernel.
That bigger nets are easier to train — and that local minima mostly go away — is an old intuition. Theories came later. The idea that had a hard time becoming mainstream is the one that contradicted every statistics textbook: a widely over-parameterized net can still generalize well.