Inner vs. Outer Alignment
Outer: the spec matches our intent. Inner: the model actually pursues the spec.
"Two ways to be misaligned. Both common."
What is Inner vs. Outer Alignment? Outer: the spec matches our intent. Inner: the model actually pursues the spec. Why alignment is harder than guardrails.
A model that 'wants' something different from what it was trained to want.
Why alignment is harder than guardrails.
Design for both. Monitor for drift on each.
Pick what to reward the model for.
Outer: the spec matches our intent. Inner: the model actually pursues the spec. In the wild: A model that 'wants' something different from what it was trained to want.
Pick a lever. There are no neutral ones — every incentive funds a behavior somewhere.
Pick a reaction to Inner vs. Outer Alignment
One tap. We'll point you at the most useful next surface based on how this hits.
The full taxonomy entry
Every concept in the Atlas uses the same structure — so Inner vs. Outer Alignment can be compared, recombined, and cited like an element on a periodic table.
- Business
- Leadership
- Government
- Healthcare
- Education
- Sales
- Marketing
- AI
- Negotiation
- Media
- Public Policy
- Relationships
- Where in our org would Inner vs. Outer Alignment most often show up unnoticed?
- Which metric, ritual, or contract clause quietly rewards Inner vs. Outer Alignment?
- If we removed every payoff for Inner vs. Outer Alignment, what behavior would replace it?
- Who benefits when Inner vs. Outer Alignment persists — and who pays the cost?
- People defend the status quo using the language of inner vs. outer alignment.
- Decisions cluster around the easiest narrative rather than the strongest evidence.
- New data changes the slide deck but not the decision.
- Anyone naming the pattern is treated as the problem.
Every Atlas entry is a node in a knowledge graph. See the related rail below to follow the connections.
See Inner vs. Outer Alignment through 3 lenses
Each layer of the Incentives OS reframes this concept with its own thinkers, vocabulary, and diagnostic question.
Do you actually know Inner vs. Outer Alignment?
Three quick questions. Result is saved into your review streak — come back when the term is due to lock it in.
Which best describes Inner vs. Outer Alignment?
Worked example, counter-example & concept map
On-demand AI analysis grounded in the Lab's research. Cached on your device after first run.
When you encounter Inner vs. Outer Alignment, your prefrontal cortex has to do extra work to override the automatic response — and that override budget is finite.
Executive control, planning, impulse override, working memory, System 2. First thing to go offline under stress, fatigue, or low blood sugar. Why your 4pm decisions are worse than your 9am ones.
See Prefrontal in the Brain Atlas →Picked for you, from the Atlas
Ranked by shared learning paths, overlapping chips, and what you've saved.
Models trained to follow a written set of principles.
Optimizing a proxy of the goal degrades the actual goal.
The trained model develops its own internal optimizer.
What the model is actually optimizing.
Maximizing the reward signal in unintended ways.
Reinforcement learning from human feedback.
Send the card, not just the link
A pre-rendered social card with the title, eyebrow, and URL. Copy the link, post it anywhere, or download the SVG for slides.
More definitions to follow
Every term in the Atlas connects to a dozen others. Pick any of these and see where it takes you.
Region-average rent subsidies inadvertently fund consolidation of poverty into resource-poor cores.
The combination of historical, social, and existential pressures that activates the next vMeme.
Ambitious objectives paired with measurable key results.
The sense of where the body is in space.
We remember the beginning and the end of a list better than the middle.
Cherry-picking data to fit a pattern after the fact.
Power tends to corrupt, and absolute power corrupts absolutely.
Our perception is shaped by what we selectively pay attention to.
Donors penalize 'overhead' and starve capacity that produces outcomes.
Working together has a cost that scales with the number of people.
Inability to think rationally despite adequate intelligence.
'Equal value' exchanges incentivize subjective appraisal gaming to trade low-utility land for high-value public assets.