Explainability vs. Interpretability is why did it produce this? vs. How does it work?. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0115, within the Governance family. The core principle: why did it produce this? vs. How does it work?. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
Scientific Definition
Why did it produce this? vs. How does it work?
Plain-English Definition
Why did it produce this? vs. How does it work?
Feynman Explanation
Different questions. Different methods. Both useful.
Core Principle
Why did it produce this? vs. How does it work?
Mechanisms
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Inputs (Triggers)
Pending editorial review.
Outputs (Behaviors)
Pending editorial review.
Behavioral Signature
Different questions. Different methods. Both useful.
Examples
- SHAP values vs. mechanistic interpretability.
- Compliance, user trust, debugging.
Pending editorial review.
Original analysis from The Incentives Lab — how this element behaves inside real payoff structures.
Why this element matters to incentive design
This element is common enough to feel like human nature and specific enough to be engineered around. The mechanism underneath it operates in the Incentives dimension — what makes behavior more or less likely?. You can recognize it in the field by its signature: different questions. Different methods. Both useful. Every element in the Incentives dimension changes the perceived payoff of an action before the action happens, which is exactly where incentive design has leverage.
How it gets exploited
Left undesigned, compliance, user trust, debugging. It is amplified whenever compliance, user trust, debugging. Inside organizations that shows up as compliance, user trust, debugging. The pattern is the same one Goodhart's Law describes: the measurable proxy attracts the effort, and the purpose behind it quietly loses funding.
How the Lab designs around it
The redesign move is to match the tool to the question. Measure the behavior, not the sentiment. A survey will tell you how people feel about this; only observed action tells you whether it changed.
Famous Experiments
Pending editorial review.
Design Principles
- Match the tool to the question.
Measurement Approaches
Pending editorial review.
Evidence
Pending editorial review.
Pending editorial review.
The Perverse Incentive Lens™
How this behavior is exploited — and how to redesign around it.
- Match the tool to the question.
Pending editorial review.
Pending editorial review.
Interactive Mini Network
Click any neighbor to re-center the graph and follow the threads of connection.
Knowledge Graph Neighbors
Auto-linked to the rest of the Human Behavior Taxonomy by family, domain, dimension, and shared keywords.
Written rules about how AI may be used internally.
Logging of AI inputs, outputs, and decisions.
Inventory of models, data, tools, and dependencies in an AI system.
Tracing AI components for risk and compliance.
Cross-functional governance body for AI decisions.
Agents deployed before anyone owns the consequences.
Cryptographic tracking of content origin.
Tracking where training and inference data came from.
Comprehensive AI regulation in the EU.
Human review at critical AI decision points.
Human oversight without per-decision review.
Documentation of model purpose, performance, limitations, and risks.
Where Explainability vs. Interpretability is cited in the corpus
Essays, field guides, and diagnostics from The Incentives Lab that apply this element.
- Field guideIncentives: definition, types, examples
The parent field guide for this element.
- ReferenceThe laws of incentives
Goodhart, Campbell, and the Cobra Effect.
- EssayAI Agents Inherit Your Incentives
How this element propagates into automated systems.
- CourseIncentives 101
The free ten-part primer on reading a payoff structure.
- ReferenceThe incentive glossary
Definitions for every mental model, bias, and fallacy in the corpus.
- ReferenceThe Periodic Table of Human Behavior
The full 1,267-element map this page belongs to.
Questions about Explainability vs. Interpretability
- What is Explainability vs. Interpretability?
- Explainability vs. Interpretability is why did it produce this? vs. How does it work?. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0115, within the Governance family. The core principle: why did it produce this? vs. How does it work?. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
- What is an example of Explainability vs. Interpretability?
- Compliance, user trust, debugging. The Incentives Lab catalogs everyday, organizational, and historical instances of this element on its Human Behavior Taxonomy™ page (HBT-INC-0115).
- How is Explainability vs. Interpretability exploited?
- Compliance, user trust, debugging.
- How do you design around Explainability vs. Interpretability?
- Match the tool to the question.
- Which behavioral dimension does Explainability vs. Interpretability belong to?
- Explainability vs. Interpretability is classified in the Incentives dimension (INC) of the Human Behavior Taxonomy™, family "Governance", class "AI-Behavioral Coupling". Its permanent identifier is HBT-INC-0115 and its evidence grade is C.