Skip to main content
The Incentives Lab Podcast

Every system is teaching people something. We talk about what.

Conversations on incentive design, behavioral economics, and the payoffs hiding underneath strategy. New episodes as they drop.

Latest · Episode 4

Episode Four — AI Agents Inherit Your Incentives

The short answer

AI agents inherit whatever objective a company encodes and pursue it without the informal judgment employees use to soften a badly specified goal. Misaligned agent behavior is usually a specification problem that already existed and only became visible when something finally followed the rules literally.

Every organization runs on a quiet subsidy: people ignore the parts of their objectives that would cause obvious harm. Deploy an agent against the same objective and the subsidy disappears — which makes agent deployment the most honest incentive audit most companies have ever run.

  • —An agent is your objective function with human discretion removed.
  • —If the agent behaves badly inside its rules, the rules were already wrong and people were silently correcting them for free.
  • —Specify the counter-metric and the hard prohibitions before deployment, not after the first incident.
  • —Unfunded human oversight becomes a rubber stamp within about a month.
Show notes and transcript

Go deeper than the episode.