A thread you can test
How Large Models Get Smaller
4 notes move from the word to a real choice at work — understand it first, then decide whether to use it.
Each note stands alone, or becomes the next step in this thread.
THE QUESTION THIS PAGE ANSWERS
ANSWER FIRSTWhat is How Large Models Get Smaller, and which AI decisions does it change?
Capabilities jump in steps once a model crosses a scale threshold — plus the academic dispute this phenomenon is still under This page keeps the related concepts, common mistakes, and practical notes in one reading thread.
First decide whether you are blocked by a definition, a choice, or verification; then choose the closest of the 4 notes below.
Start with “Emergence: Why Capabilities Appear Suddenly,” then restate the conclusion using your own task.
Do not treat every method in a topic as interchangeable. The answer changes with the input, risk, and acceptance bar.
THIS QUESTION THREAD
Put the word back inside the choice it changes.
Emergence: Why Capabilities Appear Suddenly
Capabilities jump in steps once a model crosses a scale threshold — plus the academic dispute this phenomenon is still under
Why Make Models Smaller
Three practical motives — cost, speed, on-premise deployment — and the things small models cannot do
How Distillation Works: From Teacher to Student
The five-step pipeline, soft labels, and temperature; using the six distilled models DeepSeek open-sourced alongside R1 as the sample
The Cost of Distillation: Models Are Getting More Alike
Verbal tics, formatting quirks, and identity confusion inherited wholesale; why multi-model cross-validation may be fake