Eval Drift
Evaluation suites that no longer reflect real-world conditions.
"Your benchmarks aged out of relevance."
What is Eval Drift? Evaluation suites that no longer reflect real-world conditions. Quality erosion masked by stale measurement.
Benchmark scores rising while real usage quality declines.
Quality erosion masked by stale measurement.
Periodically refresh evals against current production data.
Pick what to reward the model for.
Evaluation suites that no longer reflect real-world conditions. In the wild: Benchmark scores rising while real usage quality declines.
Pick a lever. There are no neutral ones — every incentive funds a behavior somewhere.
Pick a reaction to Eval Drift
One tap. We'll point you at the most useful next surface based on how this hits.
The full taxonomy entry
Every concept in the Atlas uses the same structure — so Eval Drift can be compared, recombined, and cited like an element on a periodic table.
- Business
- Leadership
- Government
- Healthcare
- Education
- Sales
- Marketing
- AI
- Negotiation
- Media
- Public Policy
- Relationships
- Where in our org would Eval Drift most often show up unnoticed?
- Which metric, ritual, or contract clause quietly rewards Eval Drift?
- If we removed every payoff for Eval Drift, what behavior would replace it?
- Who benefits when Eval Drift persists — and who pays the cost?
- People defend the status quo using the language of eval drift.
- Decisions cluster around the easiest narrative rather than the strongest evidence.
- New data changes the slide deck but not the decision.
- Anyone naming the pattern is treated as the problem.
Every Atlas entry is a node in a knowledge graph. See the related rail below to follow the connections.
See Eval Drift through 5 lenses
Each layer of the Incentives OS reframes this concept with its own thinkers, vocabulary, and diagnostic question.
- Layer 1Existing Framework
Which developmental stage and archetype is driving this?
- Layer 4Systems Thinking
What feedback loop is reinforcing this behavior?
- Layer 11Economics & Mechanism Design
Who pays, who is paid, and what does the price signal hide?
- Layer 15AI & Alignment
What proxy reward is the AI optimizing — and what is it ignoring?
- Layer 19Human Needs & Meaning
Which basic human need is being met — or starved — by this design?
Do you actually know Eval Drift?
Three quick questions. Result is saved into your review streak — come back when the term is due to lock it in.
Which best describes Eval Drift?
Worked example, counter-example & concept map
On-demand AI analysis grounded in the Lab's research. Cached on your device after first run.
When you encounter Eval Drift, your prefrontal cortex has to do extra work to override the automatic response — and that override budget is finite.
Executive control, planning, impulse override, working memory, System 2. First thing to go offline under stress, fatigue, or low blood sugar. Why your 4pm decisions are worse than your 9am ones.
See Prefrontal in the Brain Atlas →Picked for you, from the Atlas
Ranked by shared learning paths, overlapping chips, and what you've saved.
AI system that takes actions to achieve goals, often across tools.
Designing processes from scratch around AI capability.
The relationship between inputs and outputs changes.
How much input the model can process at once.
Underlying data distribution changes over time.
Vector representation of content for similarity and search.
Send the card, not just the link
A pre-rendered social card with the title, eyebrow, and URL. Copy the link, post it anywhere, or download the SVG for slides.
More definitions to follow
Every term in the Atlas connects to a dozen others. Pick any of these and see where it takes you.
Everything takes longer than you expect, even when you account for Hofstadter's Law.
Headcount cuts pop short-term margin but collapse morale, institutional knowledge, and execution.
Systems with weak corrective feedback drift unchecked into failure modes.
Herbert Simon's frame: agents use heuristics that work under real cognitive limits, not impossible global optima.
Humans need autonomy, competence, and relatedness.
A combination of complementary skills that produces compounding advantage greater than the sum of parts.
Productivity targets compress visits, raising misdiagnosis and burnout.
Saying it often enough until objections fade.
Treating one kind of thing as if it belonged to a different ontological category.
The same message in a different setting produces a different result.
Resolving doubt by jumping to a conclusion — any conclusion — quickly.
For a claim to be scientific, it must be possible to prove it false.