Skip to main content
HBT-INC-0222 · Dimension INC · Incentives

Prompt Injection

Malicious instructions hidden in user input or retrieved content.

Risk·AI-Behavioral Coupling·Grade C·draft· enriching…
In one paragraph

Prompt Injection is malicious instructions hidden in user input or retrieved content. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0222, within the Risk family. The core principle: malicious instructions hidden in user input or retrieved content. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.

Scientific Definition

Malicious instructions hidden in user input or retrieved content.

Plain-English Definition

Malicious instructions hidden in user input or retrieved content.

Feynman Explanation

Your agent's loyalty just got rewritten by a webpage.

Core Principle

Malicious instructions hidden in user input or retrieved content.

Mechanisms

Psychological

Pending editorial review.

Behavioral Economic

Pending editorial review.

Neurological

Pending editorial review.

Evolutionary

Pending editorial review.

Sociological

Pending editorial review.

Computational

Pending editorial review.

Systems

Pending editorial review.

Inputs (Triggers)

Pending editorial review.

Outputs (Behaviors)

Pending editorial review.

Behavioral Signature

Your agent's loyalty just got rewritten by a webpage.

Examples

Everyday
  • Indirect prompt injection through documents an agent reads.
Modern (Organizational)
  • Major security risk for agent deployments.
Historical

Pending editorial review.

Lab Commentary

Original analysis from The Incentives Lab — how this element behaves inside real payoff structures.

Why this element matters to incentive design

This is one of the elements leaders describe as a values gap. It is a payoff gap. The mechanism underneath it operates in the Incentives dimension — what makes behavior more or less likely?. You can recognize it in the field by its signature: your agent's loyalty just got rewritten by a webpage. Every element in the Incentives dimension changes the perceived payoff of an action before the action happens, which is exactly where incentive design has leverage.

How it gets exploited

Left undesigned, major security risk for agent deployments. It is amplified whenever major security risk for agent deployments. Inside organizations that shows up as major security risk for agent deployments. The pattern is the same one Goodhart's Law describes: the measurable proxy attracts the effort, and the purpose behind it quietly loses funding.

How the Lab designs around it

The redesign move is to treat all input as adversarial. Output validation. Sandboxing. Design against it the way you would design against a known failure mode — assume it will appear, and price the exploit before someone finds it.

Famous Experiments

Pending editorial review.

Design Principles

  • Treat all input as adversarial. Output validation. Sandboxing.

Measurement Approaches

Pending editorial review.

Evidence

Evidence Grade
C (A strongest → E speculative)
Replication
★★☆☆☆
Intervention Confidence
3 / 5
Consensus
Pending editorial review (HBT v1.0 auto-seed).
Limitations
Pending editorial review (HBT v1.0 auto-seed).
Open Research Questions

Pending editorial review.

Primary References

Pending editorial review.

Signature Section

The Perverse Incentive Lens™

How this behavior is exploited — and how to redesign around it.

Exploitation
Major security risk for agent deployments.
Amplifying Incentives
Major security risk for agent deployments.
Org Failure Modes
Major security risk for agent deployments.
Societal Failure Modes
Pending editorial review (HBT v1.0 auto-seed).
Ethical Considerations
Pending editorial review (HBT v1.0 auto-seed).
Redesign Strategies
Treat all input as adversarial. Output validation. Sandboxing.
Diagnostic Questions
  • Treat all input as adversarial. Output validation. Sandboxing.
Warning Signs

Pending editorial review.

Red Flags

Pending editorial review.

Intervention Playbook
Individual
Treat all input as adversarial. Output validation. Sandboxing.
Team
Pending editorial review (HBT v1.0 auto-seed).
Organization
Pending editorial review (HBT v1.0 auto-seed).
Policy
Pending editorial review (HBT v1.0 auto-seed).
AI Implications
Detection
Pending editorial review (HBT v1.0 auto-seed).
Measurement
Pending editorial review (HBT v1.0 auto-seed).
Mitigation
Pending editorial review (HBT v1.0 auto-seed).
Responsible Use
Pending editorial review (HBT v1.0 auto-seed).

Interactive Mini Network

Click any neighbor to re-center the graph and follow the threads of connection.

HBT-INC-0222 · INC
Prompt Injection
PIALAgentic LiabilityARAI Risk TieringBIBias in AI SystemsCTCounterfactual TestingDPDifferential PrivacyFMFairness MetricsFLFederated LearningJaJailbreakPMProductivity MirageRARed-Teaming AI

Knowledge Graph Neighbors

Where Prompt Injection is cited in the corpus

Questions about Prompt Injection

What is Prompt Injection?
Prompt Injection is malicious instructions hidden in user input or retrieved content. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0222, within the Risk family. The core principle: malicious instructions hidden in user input or retrieved content. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
What is an example of Prompt Injection?
Major security risk for agent deployments. The Incentives Lab catalogs everyday, organizational, and historical instances of this element on its Human Behavior Taxonomy™ page (HBT-INC-0222).
How is Prompt Injection exploited?
Major security risk for agent deployments.
How do you design around Prompt Injection?
Treat all input as adversarial. Output validation. Sandboxing.
Which behavioral dimension does Prompt Injection belong to?
Prompt Injection is classified in the Incentives dimension (INC) of the Human Behavior Taxonomy™, family "Risk", class "AI-Behavioral Coupling". Its permanent identifier is HBT-INC-0222 and its evidence grade is C.

Version History

v1.1.0 · 2026-06-28Initial auto-seed from corpus.