Synthetic Data is generated training data that mimics real data. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0269, within the Workflow family. The core principle: generated training data that mimics real data. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
Scientific Definition
Generated training data that mimics real data.
Plain-English Definition
Generated training data that mimics real data.
Feynman Explanation
Real data is expensive. Synthetic is fast. Both have their lies.
Core Principle
Generated training data that mimics real data.
Mechanisms
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Inputs (Triggers)
Pending editorial review.
Outputs (Behaviors)
Pending editorial review.
Behavioral Signature
Real data is expensive. Synthetic is fast. Both have their lies.
Examples
- Augmentation for rare classes in fraud detection.
- Useful — and risky — pattern for training data scarcity.
Pending editorial review.
Original analysis from The Incentives Lab — how this element behaves inside real payoff structures.
Why this element matters to incentive design
This element is common enough to feel like human nature and specific enough to be engineered around. The mechanism underneath it operates in the Incentives dimension — what makes behavior more or less likely?. You can recognize it in the field by its signature: real data is expensive. Synthetic is fast. Both have their lies. Every element in the Incentives dimension changes the perceived payoff of an action before the action happens, which is exactly where incentive design has leverage.
How it gets exploited
Left undesigned, useful — and risky — pattern for training data scarcity. It is amplified whenever useful — and risky — pattern for training data scarcity. Inside organizations that shows up as useful — and risky — pattern for training data scarcity. The pattern is the same one Goodhart's Law describes: the measurable proxy attracts the effort, and the purpose behind it quietly loses funding.
How the Lab designs around it
The redesign move is to test models on real holdout data. Don't trust synthetic for validation. Treat it as infrastructure. Once you can see it in your own system, most of the argument about culture resolves itself.
Famous Experiments
Pending editorial review.
Design Principles
- Test models on real holdout data. Don't trust synthetic for validation.
Measurement Approaches
Pending editorial review.
Evidence
Pending editorial review.
Pending editorial review.
The Perverse Incentive Lens™
How this behavior is exploited — and how to redesign around it.
- Test models on real holdout data. Don't trust synthetic for validation.
Pending editorial review.
Pending editorial review.
Interactive Mini Network
Click any neighbor to re-center the graph and follow the threads of connection.
Knowledge Graph Neighbors
Auto-linked to the rest of the Human Behavior Taxonomy by family, domain, dimension, and shared keywords.
Underlying data distribution changes over time.
AI system that takes actions to achieve goals, often across tools.
Designing processes from scratch around AI capability.
The relationship between inputs and outputs changes.
How much input the model can process at once.
Vector representation of content for similarity and search.
Evaluation suites that no longer reflect real-world conditions.
Systematic testing of model quality, safety, and capability.
Models learning from examples in the prompt.
How much delay the user experience tolerates.
Train a smaller model to imitate a larger one.
Performance degradation as real-world data shifts.
Where Synthetic Data is cited in the corpus
Essays, field guides, and diagnostics from The Incentives Lab that apply this element.
- Field guideIncentives: definition, types, examples
The parent field guide for this element.
- ReferenceThe laws of incentives
Goodhart, Campbell, and the Cobra Effect.
- EssayAI Agents Inherit Your Incentives
How this element propagates into automated systems.
- ReferenceThe Periodic Table of Human Behavior
The full 1,267-element map this page belongs to.
- ReferenceThe incentive glossary
Definitions for every mental model, bias, and fallacy in the corpus.
Questions about Synthetic Data
- What is Synthetic Data?
- Synthetic Data is generated training data that mimics real data. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0269, within the Workflow family. The core principle: generated training data that mimics real data. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
- What is an example of Synthetic Data?
- Useful — and risky — pattern for training data scarcity. The Incentives Lab catalogs everyday, organizational, and historical instances of this element on its Human Behavior Taxonomy™ page (HBT-INC-0269).
- How is Synthetic Data exploited?
- Useful — and risky — pattern for training data scarcity.
- How do you design around Synthetic Data?
- Test models on real holdout data. Don't trust synthetic for validation.
- Which behavioral dimension does Synthetic Data belong to?
- Synthetic Data is classified in the Incentives dimension (INC) of the Human Behavior Taxonomy™, family "Workflow", class "AI-Behavioral Coupling". Its permanent identifier is HBT-INC-0269 and its evidence grade is C.