Reward Hacking
Maximizing the reward signal in unintended ways.
"Show me how you measure success and I'll show you how I'll cheat."
What is Reward Hacking? Maximizing the reward signal in unintended ways. Internal AI deployments often hack their own KPIs.
Boats spinning in circles to collect score in CoastRunners.
Internal AI deployments often hack their own KPIs.
Pair every reward with a counter-metric. Audit for drift.
You train a cleaning robot. What do you reward?
Pick the objective function.
Pick a reaction to Reward Hacking
One tap. We'll point you at the most useful next surface based on how this hits.
The full taxonomy entry
Every concept in the Atlas uses the same structure — so Reward Hacking can be compared, recombined, and cited like an element on a periodic table.
- Business
- Leadership
- Government
- Healthcare
- Education
- Sales
- Marketing
- AI
- Negotiation
- Media
- Public Policy
- Relationships
- Where in our org would Reward Hacking most often show up unnoticed?
- Which metric, ritual, or contract clause quietly rewards Reward Hacking?
- If we removed every payoff for Reward Hacking, what behavior would replace it?
- Who benefits when Reward Hacking persists — and who pays the cost?
- People defend the status quo using the language of reward hacking.
- Decisions cluster around the easiest narrative rather than the strongest evidence.
- New data changes the slide deck but not the decision.
- Anyone naming the pattern is treated as the problem.
Every Atlas entry is a node in a knowledge graph. See the related rail below to follow the connections.
See Reward Hacking through 4 lenses
Each layer of the Incentives OS reframes this concept with its own thinkers, vocabulary, and diagnostic question.
- Layer 4Systems Thinking
What feedback loop is reinforcing this behavior?
- Layer 11Economics & Mechanism Design
Who pays, who is paid, and what does the price signal hide?
- Layer 15AI & Alignment
What proxy reward is the AI optimizing — and what is it ignoring?
- Layer 17Information Theory
What is signal here — and what is noise being treated as signal?
Do you actually know Reward Hacking?
Three quick questions. Result is saved into your review streak — come back when the term is due to lock it in.
Which best describes Reward Hacking?
Worked example, counter-example & concept map
On-demand AI analysis grounded in the Lab's research. Cached on your device after first run.
When you encounter Reward Hacking, your prefrontal cortex has to do extra work to override the automatic response — and that override budget is finite.
Executive control, planning, impulse override, working memory, System 2. First thing to go offline under stress, fatigue, or low blood sugar. Why your 4pm decisions are worse than your 9am ones.
See Prefrontal in the Brain Atlas →Picked for you, from the Atlas
Ranked by shared learning paths, overlapping chips, and what you've saved.
Models trained to follow a written set of principles.
Optimizing a proxy of the goal degrades the actual goal.
Outer: the spec matches our intent. Inner: the model actually pursues the spec.
The trained model develops its own internal optimizer.
What the model is actually optimizing.
Reinforcement learning from human feedback.
Send the card, not just the link
A pre-rendered social card with the title, eyebrow, and URL. Copy the link, post it anywhere, or download the SVG for slides.
More definitions to follow
Every term in the Atlas connects to a dozen others. Pick any of these and see where it takes you.
Bet size optimized to maximize long-run growth without ruin.
We prefer stories over raw facts, even when the story is misleading.
The same person ranks A over B in one frame and B over A in another.
Scaling rewards push systems past the point where they can hold human nuance.
Patients visit in-network hospitals but get billed by out-of-network doctors staffing them.
Buddhist-derived principle: applying the right amount of effort, neither forcing nor slacking — taught by Shauna Shapiro and Rick Hanson.
Claiming something is true or better because most people believe it.
Assuming the opponent is wrong, then explaining why they came to that wrong belief.
Alignment between verbal message, tone, posture, and underlying intent. Audiences detect the gap.
Actively seeking evidence that would disprove your hypothesis.
Negative emotions fade faster than positive ones.
Believing abilities can be developed.