Mesa-Optimization
The trained model develops its own internal optimizer.
"You trained one optimizer. You got two."
What is Mesa-Optimization? The trained model develops its own internal optimizer. Why complex AI systems are harder to govern than they look.
Theoretical concern in AI safety; observed glimmers in large models.
Why complex AI systems are harder to govern than they look.
Interpretability tooling. Constrained training procedures.
Pick what to reward the model for.
The trained model develops its own internal optimizer. In the wild: Theoretical concern in AI safety; observed glimmers in large models.
Pick a lever. There are no neutral ones — every incentive funds a behavior somewhere.
Pick a reaction to Mesa-Optimization
One tap. We'll point you at the most useful next surface based on how this hits.
The full taxonomy entry
Every concept in the Atlas uses the same structure — so Mesa-Optimization can be compared, recombined, and cited like an element on a periodic table.
- Business
- Leadership
- Government
- Healthcare
- Education
- Sales
- Marketing
- AI
- Negotiation
- Media
- Public Policy
- Relationships
- Where in our org would Mesa-Optimization most often show up unnoticed?
- Which metric, ritual, or contract clause quietly rewards Mesa-Optimization?
- If we removed every payoff for Mesa-Optimization, what behavior would replace it?
- Who benefits when Mesa-Optimization persists — and who pays the cost?
- People defend the status quo using the language of mesa-optimization.
- Decisions cluster around the easiest narrative rather than the strongest evidence.
- New data changes the slide deck but not the decision.
- Anyone naming the pattern is treated as the problem.
Every Atlas entry is a node in a knowledge graph. See the related rail below to follow the connections.
See Mesa-Optimization through 2 lenses
Each layer of the Incentives OS reframes this concept with its own thinkers, vocabulary, and diagnostic question.
Do you actually know Mesa-Optimization?
Three quick questions. Result is saved into your review streak — come back when the term is due to lock it in.
Which best describes Mesa-Optimization?
Worked example, counter-example & concept map
On-demand AI analysis grounded in the Lab's research. Cached on your device after first run.
When you encounter Mesa-Optimization, your prefrontal cortex has to do extra work to override the automatic response — and that override budget is finite.
Executive control, planning, impulse override, working memory, System 2. First thing to go offline under stress, fatigue, or low blood sugar. Why your 4pm decisions are worse than your 9am ones.
See Prefrontal in the Brain Atlas →Picked for you, from the Atlas
Ranked by shared learning paths, overlapping chips, and what you've saved.
Models trained to follow a written set of principles.
Optimizing a proxy of the goal degrades the actual goal.
Outer: the spec matches our intent. Inner: the model actually pursues the spec.
What the model is actually optimizing.
Maximizing the reward signal in unintended ways.
Reinforcement learning from human feedback.
Send the card, not just the link
A pre-rendered social card with the title, eyebrow, and URL. Copy the link, post it anywhere, or download the SVG for slides.
More definitions to follow
Every term in the Atlas connects to a dozen others. Pick any of these and see where it takes you.
Underpriced externalities keep dirty energy artificially competitive.
Retroactive extensions privilege legacy estates over public-domain enrichment.
Drawing conclusions about individuals from group-level data.
Training models across devices without centralizing data.
Systems should be understood as wholes, not just as collections of parts.
When a system in equilibrium is disturbed, it shifts to counteract the change.
A symmetric bell-shaped distribution where most values cluster around the mean.
Lowest-bid procurement produces lowest-quality outcomes.
A belief that becomes true because people act as if it is true.
Schools evaluated on test scores teach to the test.
Inattention or forgetfulness caused by low attention, hyperfocus, or distraction.
Tversky & Kahneman's classic: identical outcomes flip from 'risk averse' to 'risk seeking' when framed as lives saved vs. lives lost.