Scientific Definition
Bypassing model safety constraints.
Plain-English Definition
Bypassing model safety constraints.
Feynman Explanation
Every safety system has a creative attack surface.
Core Principle
Bypassing model safety constraints.
Mechanisms
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Inputs (Triggers)
Pending editorial review.
Outputs (Behaviors)
Pending editorial review.
Behavioral Signature
Every safety system has a creative attack surface.
Examples
- Role-playing exploits that bypass content policies.
- Risk in user-facing AI products.
Pending editorial review.
Famous Experiments
Pending editorial review.
Design Principles
- Defense in depth. Continuous red-teaming. Don't rely on model alignment alone.
Measurement Approaches
Pending editorial review.
Evidence
Pending editorial review.
Pending editorial review.
The Perverse Incentive Lens™
How this behavior is exploited — and how to redesign around it.
- Defense in depth. Continuous red-teaming. Don't rely on model alignment alone.
Pending editorial review.
Pending editorial review.
Interactive Mini Network
Click any neighbor to re-center the graph and follow the threads of connection.
Knowledge Graph Neighbors
Auto-linked to the rest of the Human Behavior Taxonomy by family, domain, dimension, and shared keywords.
When the agent acts, who's responsible?
Categorizing AI use cases by risk level.
Systematic skew in model behavior across groups.
Testing model behavior on hypothetical alternate inputs.
Adding noise to data to protect individual privacy.
Quantitative measures of model behavior across groups.
Training models across devices without centralizing data.
Individual speed gains hide collective quality decline.
Malicious instructions hidden in user input or retrieved content.
Adversarial testing of AI systems.
Foundational skills erode through AI offloading.
Concentration risk on a single AI provider.