Skip to main content
The Incentives Lab
All Layers / Layer 15
Incentives OS · Layer 15

AI & Alignment

Agentic systems, reward hacking, specification gaming, goal misgeneralization. AI has perverse incentives too.

What proxy reward is the AI optimizing — and what is it ignoring?

Canonical thinkers

  • Stuart Russell
  • Paul Christiano
  • Yoshua Bengio
  • Anthropic alignment team

Seed concepts

Underlined seeds link to their full glossary entry. Plain seeds are pending a definition page.

Typed connections

Full graph →

How this layer reinforces, counteracts, or depends on the rest of the system.

  • Depends on·Layer
    This layer depends on Game Theory

    Alignment is a multi-agent game; AI safety reduces to mechanism design once capability is high enough.

  • Depends on·Layer
    This layer depends on Information Theory

    Modern AI is information compression at scale — Shannon's bounds shape what alignment can do.

Elements in this layer

547 HBEs
INC15-Minute Appointments

Productivity targets compress visits, raising misdiagnosis and burnout.

COGAbsent-Mindedness

Inattention or forgetfulness caused by low attention, hyperfocus, or distraction.

COGAbstractions

Deriving general rules from specific examples; the leap from instance to concept.

INCAcceptable Use Policy (AI)

Written rules about how AI may be used internally.

INCAcquisition Earn-Outs

Earn-outs designed to retain founders often demotivate the team they bought.

COGAd Hominem

Attacking the person rather than the argument.

INCAd-Supported Business Models

Free products monetize attention, structurally aligning incentives against user time well spent.

COGAdaptive Bias

The brain evolved to reason adaptively, not always truthfully, to reduce the cost of errors.

COGAdditive Bias

We solve problems by adding, even when subtracting would be better.

INCAdjunctification

Tenure-track jobs replaced by low-paid adjuncts, lowering cost and quality.

INCAdoption Curve (AI)

Innovators → early adopters → majority → laggards, AI-specific.

INCAesthetic-First Fitness

Training for appearance can crowd out mobility, longevity, and mental health.

COGAffirming a Disjunct

Assuming that if one option is true, another must be false, when both can be true.

COGAffirming the Consequent

If A then B; B happened; therefore A.

INCAgent

AI system that takes actions to achieve goals, often across tools.

COGAgent Detection

Presuming a purposeful actor behind events that may have no actor at all.

INCAgentic Liability

When the agent acts, who's responsible?

INCAI Agent Liability Vacuum

Autonomous agents deployed before liability frameworks exist.

INCAI as Coach vs. AI as Replacement

Augmentation strategy vs. substitution strategy.

INCAI Audit Trail

Logging of AI inputs, outputs, and decisions.

INCAI Bill of Materials

Inventory of models, data, tools, and dependencies in an AI system.

INCAI Bill of Materials Auditability

Tracing AI components for risk and compliance.

INCAI Council / Committee

Cross-functional governance body for AI decisions.

INCAI Governance Vacuum

Agents deployed before anyone owns the consequences.

INCAI Maturity Model

Stages of organizational AI capability.

INCAI Productivity Mirage

Individual productivity gains hide collective output degradation.

INCAI Productivity Paradox

AI investment outpacing measurable productivity gains.

INCAI Risk Tiering

Categorizing AI use cases by risk level.

INCAI Strategy as Capability Strategy

AI strategy = decisions about which capabilities to build and where.

INCAI Talent Premium

Scarce AI talent commands market-distorting compensation.

INCAI-First Workflow Design

Designing processes from scratch around AI capability.

INCAI-Generated Content Spam Loop

AI generates content; AI scrapes content; AI trains on its own output.

INCAI-Native vs. AI-Augmented Teams

Teams built around AI from day one vs. teams adding it to existing workflows.

COGAlder's Razor

What cannot be settled by experiment is not worth debating.

INCAlgorithmic Aversion

Discounting algorithmic advice even when superior.

COGAlgorithms

A finite set of well-defined instructions for solving a problem or performing a computation.

COGAll Models Are Wrong

Every model simplifies reality; some are still useful.

COGAllegiance Bias

Researchers favor conclusions aligned with their school, team, or sponsor.

COGAmbiguity Aversion

We prefer known risks to unknown ones, even when the unknown is better.

COGAnchoring (NLP)

Pairing a sensory cue with a desired internal state until the cue reliably evokes the state.

COGAnecdotal Evidence

A single story used as proof of a general claim.

COGAnecdotal Fallacy

Using personal stories or isolated examples instead of evidence.

BIOAnterior Cingulate Cortex

Brain region tracking conflict, error, and effort.

INCAnthropomorphism

Treating AI as more humanlike than it is.

BIOAnthropomorphism (AI)

We treat AI as more humanlike than it is.

SYSAntifragility

Systems that gain from disorder.

COGAppeal to Authority

It's true because authority says so.

COGAppeal to Common Sense

Common sense says so, therefore it's true.

COGAppeal to Consequences

Believing something is true because of the consequences of its truth.

COGAppeal to Emotion

Substituting feeling for argument.

COGAppeal to Fear

Using fear instead of evidence.

COGAppeal to Ignorance

Claiming something is true because it hasn't been proven false (or vice versa).

COGAppeal to Majority

Claiming something is true or better because most people believe it.

COGAppeal to Nature

If it's 'natural,' it's good.

COGAppeal to Novelty

It's better because it's newer.

COGAppeal to Pity

Asking for a conclusion based on sympathy rather than logic.

COGAppeal to Probability

Assuming something is true because it is probable or possible.

COGAppeal to Spite

Rejecting an argument because of dislike for the source or beneficiary.

COGAppeal to Tradition

It's right because it's how we've always done it.

COGArgument from Authority

Using an authority's opinion as evidence, regardless of its merits.

Showing the first 60 of 547. Full graph view coming in Phase 2.