Model Distillation
Train a smaller model to imitate a larger one.
"Compress the wisdom. Keep the speed."
What is Model Distillation? Train a smaller model to imitate a larger one. Cost and latency optimization.
Edge deployments of distilled LLMs.
Cost and latency optimization.
Distillation often beats fine-tuning at scale.
Pick what to reward the model for.
Train a smaller model to imitate a larger one. In the wild: Edge deployments of distilled LLMs.
Pick a lever. There are no neutral ones — every incentive funds a behavior somewhere.
Pick a reaction to Model Distillation
One tap. We'll point you at the most useful next surface based on how this hits.
The full taxonomy entry
Every concept in the Atlas uses the same structure — so Model Distillation can be compared, recombined, and cited like an element on a periodic table.
- Business
- Leadership
- Government
- Healthcare
- Education
- Sales
- Marketing
- AI
- Negotiation
- Media
- Public Policy
- Relationships
- Where in our org would Model Distillation most often show up unnoticed?
- Which metric, ritual, or contract clause quietly rewards Model Distillation?
- If we removed every payoff for Model Distillation, what behavior would replace it?
- Who benefits when Model Distillation persists — and who pays the cost?
- People defend the status quo using the language of model distillation.
- Decisions cluster around the easiest narrative rather than the strongest evidence.
- New data changes the slide deck but not the decision.
- Anyone naming the pattern is treated as the problem.
Every Atlas entry is a node in a knowledge graph. See the related rail below to follow the connections.
See Model Distillation through 4 lenses
Each layer of the Incentives OS reframes this concept with its own thinkers, vocabulary, and diagnostic question.
- Layer 1Existing Framework
Which developmental stage and archetype is driving this?
- Layer 11Economics & Mechanism Design
Who pays, who is paid, and what does the price signal hide?
- Layer 15AI & Alignment
What proxy reward is the AI optimizing — and what is it ignoring?
- Layer 19Human Needs & Meaning
Which basic human need is being met — or starved — by this design?
Do you actually know Model Distillation?
Three quick questions. Result is saved into your review streak — come back when the term is due to lock it in.
Which best describes Model Distillation?
Worked example, counter-example & concept map
On-demand AI analysis grounded in the Lab's research. Cached on your device after first run.
When you encounter Model Distillation, your prefrontal cortex has to do extra work to override the automatic response — and that override budget is finite.
Executive control, planning, impulse override, working memory, System 2. First thing to go offline under stress, fatigue, or low blood sugar. Why your 4pm decisions are worse than your 9am ones.
See Prefrontal in the Brain Atlas →Picked for you, from the Atlas
Ranked by shared learning paths, overlapping chips, and what you've saved.
AI system that takes actions to achieve goals, often across tools.
Designing processes from scratch around AI capability.
The relationship between inputs and outputs changes.
How much input the model can process at once.
Underlying data distribution changes over time.
Vector representation of content for similarity and search.
Send the card, not just the link
A pre-rendered social card with the title, eyebrow, and URL. Copy the link, post it anywhere, or download the SVG for slides.
More definitions to follow
Every term in the Atlas connects to a dozen others. Pick any of these and see where it takes you.
People change behavior when they know they are being observed.
Bet size optimized to maximize long-run growth without ruin.
A state where no player benefits from changing strategy unilaterally.
Imagine the failure beforehand; investigate it afterward.
Things feel more valuable when supply or time is limited.
Patients visit in-network hospitals but get billed by out-of-network doctors staffing them.
Buddhist-derived principle: applying the right amount of effort, neither forcing nor slacking — taught by Shauna Shapiro and Rick Hanson.
If it's 'natural,' it's good.
Whoever makes the positive claim carries the obligation to support it. Absence of disproof isn't proof.
Testing hypotheses only by looking for confirming evidence.
How much we value the present over the future.
Quantitative measures of model behavior across groups.