Skip to main content
The Incentives Lab Podcast · Episode 2

Episode Two — Why Good People Game Good Metrics

Aaron Bare · 18 min · September 16, 2026
The short answer

Good people game good metrics because a proxy only tracks the outcome while nobody is optimizing the proxy directly. Attach a reward and you create a population paid to find the cheapest path to the number, which almost never runs through the outcome.

In this episode

Goodhart's Law in practice: how a metric that was honest for years turns dishonest the moment it carries a bonus, three field cases where the dashboard improved while the business got worse, and the paired-metric fix that costs nothing to implement.

Key takeaways

  • Metric gaming is a design competence, not an ethics failure — your best people find the shortest legal path fastest.
  • Every single-number target is one shortcut away from breaking the correlation that justified it.
  • Pair each target with a counter-metric that only degrades when the shortcut is taken.
  • Move control of the metric's definition downstream of the reward, or definition drift will do the gaming quietly.

Transcript

Aaron Bare: I want to start with a sentence that gets me in trouble in board rooms. Nobody in your company is gaming your metrics. They are succeeding at them. The gaming framing makes it a character problem, and once it is a character problem the fix becomes a values statement, and values statements have never once beaten a bonus formula.

Aaron Bare: Charles Goodhart wrote the line everyone quotes in a monetary policy paper in 1975. When a measure becomes a target, it ceases to be a good measure. He was describing central banking. It survived because everyone who has ever managed anything eventually watches it happen to them personally.

Aaron Bare: Here is the mechanism underneath the quote. You pick a metric because it correlates with something you actually care about and is easier to see. Time to close correlates with resolution. Time to fill correlates with staffing health. That correlation was real. It was real because nobody was trying to move the number on purpose. The second you attach money to it, you have hired a group of intelligent people whose job is now to break that correlation as efficiently as possible.

Aaron Bare: Case one. A support organization ties bonuses to average time to close. Closure time drops by a third in two quarters. Everybody is thrilled. The number nobody was reporting upward was reopen rate, and reopen rate went up by more than half. Tickets were being closed, not resolved. Same customer, same problem, now generating three tickets instead of one. Total cost per resolved issue went up while every dashboard in the building showed improvement.

Aaron Bare: Case two. Talent team measured on time to fill. The metric improves sharply, which is exactly what you would expect from competent people. Ninety day attrition roughly doubles. The fastest way to fill a role is to lower the bar on the final interview, and the cost of that decision lands one quarter later, on a different team's ledger. That is not just a measurement problem. That is a temporal problem — reward arrives in month one, cost arrives in month four, and by then the person who made the call has a different set of priorities.

Aaron Bare: Case three is the one that should make you uncomfortable. Operations group ties site bonuses to reported incident counts. Reported incidents fall. Severity of the incidents that do get reported climbs. That divergence is the signature of a reporting problem, not a safety improvement. They took a data collection system and turned it into a liability for the exact people generating the data.

Aaron Bare: Notice what is common across all three. In no case did anybody break a rule. In no case would you fire the person. And in all three cases, the harder the team worked, the worse the underlying outcome got. That is what a well-intentioned metric with no counterweight does to a competent organization.

Aaron Bare: So the fix. Two structural edits, neither of which requires more oversight, more dashboards, or a single additional meeting. First, pair the metric. Closure time with reopen rate. Time to fill with ninety day survival. Incident count with severity distribution and near miss reports. A single number can be gamed cheaply. A pair usually cannot, because the same shortcut moves them in opposite directions, and the person taking the shortcut has to explain the second number.

Aaron Bare: Second, and this one gets skipped constantly. Move definition authority downstream of the reward. If the team being measured also decides what counts as closed, what counts as filled, what counts as an incident, then definition drift does the gaming for them and nobody breaks a rule at all. The definition should live with whoever inherits the consequence.

Aaron Bare: One more thing before we close. When you make this change, do not frame it as a control. Frame it as a correction. The message is not we caught you. The message is we asked for the wrong thing and you delivered it. That framing is not diplomacy. It is accurate, and accuracy is what keeps the next metric problem from being hidden from you.

Aaron Bare: If you want the written version with the three cases in more detail, it is in the essay Goodhart's Law in the Real World in the show notes. And if you want to know which of your own metrics is currently paying for the wrong behavior, the diagnostic linked below takes about ten minutes. Next episode we go one level down, into the four layers every incentive system runs on at the same time.

Show notes and references

Host

Aaron Bare

Aaron Bare is a strategist, Wall Street Journal-bestselling author, and the founder of The Incentives Lab. He writes and advises on incentive design inside organizations — why culture is the residue of what a company rewards, how KPIs quietly go perverse, and how AI systems inherit the incentives their designers set.

All work by Aaron Bare