Episode Four — AI Agents Inherit Your Incentives
AI agents inherit whatever objective a company encodes and pursue it without the informal judgment employees use to soften a badly specified goal. Misaligned agent behavior is usually a specification problem that already existed and only became visible when something finally followed the rules literally.
In this episode
Every organization runs on a quiet subsidy: people ignore the parts of their objectives that would cause obvious harm. Deploy an agent against the same objective and the subsidy disappears — which makes agent deployment the most honest incentive audit most companies have ever run.
Key takeaways
- —An agent is your objective function with human discretion removed.
- —If the agent behaves badly inside its rules, the rules were already wrong and people were silently correcting them for free.
- —Specify the counter-metric and the hard prohibitions before deployment, not after the first incident.
- —Unfunded human oversight becomes a rubber stamp within about a month.
Transcript
Aaron Bare: There is a subsidy running inside your company that has never appeared on a budget line. People ignore the parts of their objectives that would obviously cause harm. The quota says close deals, and the rep quietly declines the account they know churns in four months. The metric says resolve tickets, and the agent escalates the one that needs a human. None of that judgment is written down anywhere. It is donated labor, and it is holding your operating model together.
Aaron Bare: Now deploy a software agent against the exact same objective. The subsidy disappears in an afternoon. The agent has no idea which parts of the goal were meant literally, because nobody ever had to say. Nobody had to say, because for twenty years the only entities reading the objective were people who had absorbed the unwritten half by watching what happened to colleagues.
Aaron Bare: So when a leader tells me the agent is behaving strangely, my first question is not about the model. It is: what does the objective literally say, and who was quietly refusing to follow it?
Aaron Bare: This makes agent deployment an unusually honest diagnostic. It is the cheapest incentive audit available, because a specification error that has been latent for years finally gets executed faithfully in week one. You are not watching a machine malfunction. You are watching your instructions run without a filter for the first time.
Aaron Bare: The work before deployment is not prompt engineering. It is writing down the parts of the objective that were previously carried by judgment. Six questions. What outcome is this metric a proxy for, in one sentence. What is the cheapest way to move the metric while damaging that outcome. What counter-metric would that shortcut degrade, and is the agent measured on it. Where does the cost of an agent decision land, and does anything report it back. What is the agent forbidden to do even when it would improve the number. And who owns the metric's definition, and are they downstream of the agent.
Aaron Bare: That last one is the same rule from episode two, and it matters more here, not less. A human team that quietly redefines what counts moves slowly. An agent optimizing against a definition it effectively controls moves at machine speed.
Aaron Bare: Then there is the part almost every deployment gets wrong, which is that the humans around the agent have incentives too. If a manager's number depends on the throughput the agent produces, that manager is not going to be the person who raises the quality problem. That is not cynicism, that is arithmetic. And if reviewing agent output is unpaid work stacked on top of an existing quota, review becomes a rubber stamp inside of a month. I have never seen an exception.
Aaron Bare: So budget the oversight explicitly. Give it hours, give it a number, give it standing. Unfunded oversight is not oversight. It is a control that exists on an org chart and nowhere else. And here is the tell: if your human reviewers are overturning almost nothing, that is not evidence the agent is good. That is evidence review has become a formality.
Aaron Bare: The sequence I recommend is model the payoff structure first, simulate what the objective would produce, then deploy. That is what we mean by an organizational digital twin — you run the incentive on a model of the org before you run it on the org. Most of the failures I have seen would have shown up in a two week specification pass, and not because anybody inspected the model. Because somebody finally wrote down what the metric was supposed to protect.
Aaron Bare: One closing thought. Everyone is asking whether we can align AI with human values. That is a real question and it is above my pay grade. The question that is not above my pay grade, and that most companies are skipping, is whether your organization is aligned with its own stated values before you automate it. Because whatever gap exists between what you say and what you pay for, an agent will find it, scale it, and put it in a dashboard.
Aaron Bare: Essay version, the digital twin page, and the Goodhart episode are all in the show notes. If you want us to run the specification pass with you, work with us is linked as well.
Show notes and references
Aaron Bare
Aaron Bare is a strategist, Wall Street Journal-bestselling author, and the founder of The Incentives Lab. He writes and advises on incentive design inside organizations — why culture is the residue of what a company rewards, how KPIs quietly go perverse, and how AI systems inherit the incentives their designers set.
All work by Aaron BareMore episodes
Episode One — Why Incentives Beat Strategy
The opening conversation of The Incentives Lab Podcast: why every strategy deck loses to the incentive system underneath it, how perverse incentives quietly form inside good organizations, and what leaders can redesign this quarter.
Episode Two — Why Good People Game Good Metrics
Goodhart's Law in practice: how a metric that was honest for years turns dishonest the moment it carries a bonus, three field cases where the dashboard improved while the business got worse, and the paired-metric fix that costs nothing to implement.
Episode Three — The Four Layers of Every Incentive System
Economic, status, safety, and temporal. Every organization runs all four at once, and the dysfunction leaders complain about almost always lives in the disagreement between them rather than inside any single layer.