Jailbreak is bypassing model safety constraints. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0165, within the Risk family. The core principle: bypassing model safety constraints. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
Scientific Definition
Bypassing model safety constraints.
Plain-English Definition
Bypassing model safety constraints.
Feynman Explanation
Every safety system has a creative attack surface.
Core Principle
Bypassing model safety constraints.
Mechanisms
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Inputs (Triggers)
Pending editorial review.
Outputs (Behaviors)
Pending editorial review.
Behavioral Signature
Every safety system has a creative attack surface.
Examples
- Role-playing exploits that bypass content policies.
- Risk in user-facing AI products.
Pending editorial review.
Original analysis from The Incentives Lab — how this element behaves inside real payoff structures.
Why this element matters to incentive design
Executives usually notice this element only after it has cost something. By then it looks like a one-off. It is not. The mechanism underneath it operates in the Incentives dimension — what makes behavior more or less likely?. You can recognize it in the field by its signature: every safety system has a creative attack surface. Every element in the Incentives dimension changes the perceived payoff of an action before the action happens, which is exactly where incentive design has leverage.
How it gets exploited
Left undesigned, risk in user-facing AI products. It is amplified whenever risk in user-facing AI products. Inside organizations that shows up as risk in user-facing AI products. The pattern is the same one Goodhart's Law describes: the measurable proxy attracts the effort, and the purpose behind it quietly loses funding.
How the Lab designs around it
The redesign move is to defense in depth. Continuous red-teaming. Don't rely on model alignment alone. Measure the behavior, not the sentiment. A survey will tell you how people feel about this; only observed action tells you whether it changed.
Famous Experiments
Pending editorial review.
Design Principles
- Defense in depth. Continuous red-teaming. Don't rely on model alignment alone.
Measurement Approaches
Pending editorial review.
Evidence
Pending editorial review.
Pending editorial review.
The Perverse Incentive Lens™
How this behavior is exploited — and how to redesign around it.
- Defense in depth. Continuous red-teaming. Don't rely on model alignment alone.
Pending editorial review.
Pending editorial review.
Interactive Mini Network
Click any neighbor to re-center the graph and follow the threads of connection.
Knowledge Graph Neighbors
Auto-linked to the rest of the Human Behavior Taxonomy by family, domain, dimension, and shared keywords.
When the agent acts, who's responsible?
Categorizing AI use cases by risk level.
Systematic skew in model behavior across groups.
Testing model behavior on hypothetical alternate inputs.
Adding noise to data to protect individual privacy.
Quantitative measures of model behavior across groups.
Training models across devices without centralizing data.
Individual speed gains hide collective quality decline.
Malicious instructions hidden in user input or retrieved content.
Adversarial testing of AI systems.
Foundational skills erode through AI offloading.
Concentration risk on a single AI provider.
Where Jailbreak is cited in the corpus
Essays, field guides, and diagnostics from The Incentives Lab that apply this element.
- Field guideIncentives: definition, types, examples
The parent field guide for this element.
- ReferenceThe laws of incentives
Goodhart, Campbell, and the Cobra Effect.
- EssayAI Agents Inherit Your Incentives
How this element propagates into automated systems.
- EssayIncentives Under Crisis
How this element behaves under pressure.
- ReferenceThe incentive glossary
Definitions for every mental model, bias, and fallacy in the corpus.
- CourseIncentives 101
The free ten-part primer on reading a payoff structure.
- ReferenceThe Periodic Table of Human Behavior
The full 1,267-element map this page belongs to.
Questions about Jailbreak
- What is Jailbreak?
- Jailbreak is bypassing model safety constraints. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0165, within the Risk family. The core principle: bypassing model safety constraints. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
- What is an example of Jailbreak?
- Risk in user-facing AI products. The Incentives Lab catalogs everyday, organizational, and historical instances of this element on its Human Behavior Taxonomy™ page (HBT-INC-0165).
- How is Jailbreak exploited?
- Risk in user-facing AI products.
- How do you design around Jailbreak?
- Defense in depth. Continuous red-teaming. Don't rely on model alignment alone.
- Which behavioral dimension does Jailbreak belong to?
- Jailbreak is classified in the Incentives dimension (INC) of the Human Behavior Taxonomy™, family "Risk", class "AI-Behavioral Coupling". Its permanent identifier is HBT-INC-0165 and its evidence grade is C.