Notes
Papers, posts, and other bullet points worth keeping.
FLOPs were intelligence; parameters were knowledge!
- 1 of 2048
- experts per token
- <3B
- activated params
- 1.6T
- total params
Prompted by Jie Tang’s history of scaling laws, Liam Fedus went back to the Switch Transformer era.
In 2020, they explored the limits of sparsity by routing each token to only 1 out of 2048 experts — in retrospect, a bold choice. The model had fewer than 3B activated parameters, but 1.6T total parameters, comparable to today’s frontier models.
The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE.
The lesson: the optimal tokens-per-parameter ratio is highly task-dependent.