Skip to main content
HBT-INC-0159 · Dimension INC · Incentives

Inner vs. Outer Alignment

Outer: the spec matches our intent. Inner: the model actually pursues the spec.

Alignment·AI-Behavioral Coupling·Grade C·draft· enriching…
In one paragraph

Inner vs. Outer Alignment is outer: the spec matches our intent. Inner: the model actually pursues the spec. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0159, within the Alignment family. The core principle: outer: the spec matches our intent. Inner: the model actually pursues the spec. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.

Scientific Definition

Outer: the spec matches our intent. Inner: the model actually pursues the spec.

Plain-English Definition

Outer: the spec matches our intent. Inner: the model actually pursues the spec.

Feynman Explanation

Two ways to be misaligned. Both common.

Core Principle

Outer: the spec matches our intent. Inner: the model actually pursues the spec.

Mechanisms

Psychological

Pending editorial review.

Behavioral Economic

Pending editorial review.

Neurological

Pending editorial review.

Evolutionary

Pending editorial review.

Sociological

Pending editorial review.

Computational

Pending editorial review.

Systems

Pending editorial review.

Inputs (Triggers)

Pending editorial review.

Outputs (Behaviors)

Pending editorial review.

Behavioral Signature

Two ways to be misaligned. Both common.

Examples

Everyday
  • A model that 'wants' something different from what it was trained to want.
Modern (Organizational)
  • Why alignment is harder than guardrails.
Historical

Pending editorial review.

Lab Commentary

Original analysis from The Incentives Lab — how this element behaves inside real payoff structures.

Why this element matters to incentive design

Most organizations meet this element as a personnel problem. It is not one. The mechanism underneath it operates in the Incentives dimension — what makes behavior more or less likely?. You can recognize it in the field by its signature: two ways to be misaligned. Both common. Every element in the Incentives dimension changes the perceived payoff of an action before the action happens, which is exactly where incentive design has leverage.

How it gets exploited

Left undesigned, why alignment is harder than guardrails. It is amplified whenever why alignment is harder than guardrails. Inside organizations that shows up as why alignment is harder than guardrails. The pattern is the same one Goodhart's Law describes: the measurable proxy attracts the effort, and the purpose behind it quietly loses funding.

How the Lab designs around it

The redesign move is to design for both. Monitor for drift on each. The leverage is not in explaining the behavior to people. It is in changing what the behavior earns.

Famous Experiments

Pending editorial review.

Design Principles

  • Design for both. Monitor for drift on each.

Measurement Approaches

Pending editorial review.

Evidence

Evidence Grade
C (A strongest → E speculative)
Replication
★★☆☆☆
Intervention Confidence
3 / 5
Consensus
Pending editorial review (HBT v1.0 auto-seed).
Limitations
Pending editorial review (HBT v1.0 auto-seed).
Open Research Questions

Pending editorial review.

Primary References

Pending editorial review.

Signature Section

The Perverse Incentive Lens™

How this behavior is exploited — and how to redesign around it.

Exploitation
Why alignment is harder than guardrails.
Amplifying Incentives
Why alignment is harder than guardrails.
Org Failure Modes
Why alignment is harder than guardrails.
Societal Failure Modes
Pending editorial review (HBT v1.0 auto-seed).
Ethical Considerations
Pending editorial review (HBT v1.0 auto-seed).
Redesign Strategies
Design for both. Monitor for drift on each.
Diagnostic Questions
  • Design for both. Monitor for drift on each.
Warning Signs

Pending editorial review.

Red Flags

Pending editorial review.

Intervention Playbook
Individual
Design for both. Monitor for drift on each.
Team
Pending editorial review (HBT v1.0 auto-seed).
Organization
Pending editorial review (HBT v1.0 auto-seed).
Policy
Pending editorial review (HBT v1.0 auto-seed).
AI Implications
Detection
Pending editorial review (HBT v1.0 auto-seed).
Measurement
Pending editorial review (HBT v1.0 auto-seed).
Mitigation
Pending editorial review (HBT v1.0 auto-seed).
Responsible Use
Pending editorial review (HBT v1.0 auto-seed).

Interactive Mini Network

Click any neighbor to re-center the graph and follow the threads of connection.

HBT-INC-0159 · INC
Inner vs. Outer Alignment
IVCAConstitutional AIGLGoodhart's Law (AI f…MeMesa-OptimizationOFObjective FunctionRHReward HackingRLRLHFSGSpecification GamingAUAcceptable Use Polic…ACAdoption Curve (AI)AgAgent

Knowledge Graph Neighbors

Where Inner vs. Outer Alignment is cited in the corpus

Questions about Inner vs. Outer Alignment

What is Inner vs. Outer Alignment?
Inner vs. Outer Alignment is outer: the spec matches our intent. Inner: the model actually pursues the spec. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0159, within the Alignment family. The core principle: outer: the spec matches our intent. Inner: the model actually pursues the spec. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
What is an example of Inner vs. Outer Alignment?
Why alignment is harder than guardrails. The Incentives Lab catalogs everyday, organizational, and historical instances of this element on its Human Behavior Taxonomy™ page (HBT-INC-0159).
How is Inner vs. Outer Alignment exploited?
Why alignment is harder than guardrails.
How do you design around Inner vs. Outer Alignment?
Design for both. Monitor for drift on each.
Which behavioral dimension does Inner vs. Outer Alignment belong to?
Inner vs. Outer Alignment is classified in the Incentives dimension (INC) of the Human Behavior Taxonomy™, family "Alignment", class "AI-Behavioral Coupling". Its permanent identifier is HBT-INC-0159 and its evidence grade is C.

Version History

v1.1.0 · 2026-06-28Initial auto-seed from corpus.