Skip to main content
The Incentives Lab
HBT-INC-0165 · Dimension INC · Incentives

Jailbreak

Bypassing model safety constraints.

Risk·AI-Behavioral Coupling·Grade C·draft· enriching…

Scientific Definition

Bypassing model safety constraints.

Plain-English Definition

Bypassing model safety constraints.

Feynman Explanation

Every safety system has a creative attack surface.

Core Principle

Bypassing model safety constraints.

Mechanisms

Psychological

Pending editorial review.

Behavioral Economic

Pending editorial review.

Neurological

Pending editorial review.

Evolutionary

Pending editorial review.

Sociological

Pending editorial review.

Computational

Pending editorial review.

Systems

Pending editorial review.

Inputs (Triggers)

Pending editorial review.

Outputs (Behaviors)

Pending editorial review.

Behavioral Signature

Every safety system has a creative attack surface.

Examples

Everyday
  • Role-playing exploits that bypass content policies.
Modern (Organizational)
  • Risk in user-facing AI products.
Historical

Pending editorial review.

Famous Experiments

Pending editorial review.

Design Principles

  • Defense in depth. Continuous red-teaming. Don't rely on model alignment alone.

Measurement Approaches

Pending editorial review.

Evidence

Evidence Grade
C (A strongest → E speculative)
Replication
★★☆☆☆
Intervention Confidence
3 / 5
Consensus
Pending editorial review (HBT v1.0 auto-seed).
Limitations
Pending editorial review (HBT v1.0 auto-seed).
Open Research Questions

Pending editorial review.

Primary References

Pending editorial review.

Signature Section

The Perverse Incentive Lens™

How this behavior is exploited — and how to redesign around it.

Exploitation
Risk in user-facing AI products.
Amplifying Incentives
Risk in user-facing AI products.
Org Failure Modes
Risk in user-facing AI products.
Societal Failure Modes
Pending editorial review (HBT v1.0 auto-seed).
Ethical Considerations
Pending editorial review (HBT v1.0 auto-seed).
Redesign Strategies
Defense in depth. Continuous red-teaming. Don't rely on model alignment alone.
Diagnostic Questions
  • Defense in depth. Continuous red-teaming. Don't rely on model alignment alone.
Warning Signs

Pending editorial review.

Red Flags

Pending editorial review.

Intervention Playbook
Individual
Defense in depth. Continuous red-teaming. Don't rely on model alignment alone.
Team
Pending editorial review (HBT v1.0 auto-seed).
Organization
Pending editorial review (HBT v1.0 auto-seed).
Policy
Pending editorial review (HBT v1.0 auto-seed).
AI Implications
Detection
Pending editorial review (HBT v1.0 auto-seed).
Measurement
Pending editorial review (HBT v1.0 auto-seed).
Mitigation
Pending editorial review (HBT v1.0 auto-seed).
Responsible Use
Pending editorial review (HBT v1.0 auto-seed).

Interactive Mini Network

Click any neighbor to re-center the graph and follow the threads of connection.

HBT-INC-0165 · INC
Jailbreak
JaALAgentic LiabilityARAI Risk TieringBIBias in AI SystemsCTCounterfactual TestingDPDifferential PrivacyFMFairness MetricsFLFederated LearningPMProductivity MiragePIPrompt InjectionRARed-Teaming AI

Knowledge Graph Neighbors

Version History

v1.1.0 · 2026-06-28Initial auto-seed from corpus.