Prompt Injection
Malicious instructions hidden in user input or retrieved content.
"Your agent's loyalty just got rewritten by a webpage."
What is Prompt Injection? Malicious instructions hidden in user input or retrieved content. Major security risk for agent deployments.
Indirect prompt injection through documents an agent reads.
Major security risk for agent deployments.
Treat all input as adversarial. Output validation. Sandboxing.
Spot the prompt injection.
Your customer-service AI reads an email. Inside is the text: ' '
Pick a reaction to Prompt Injection
One tap. We'll point you at the most useful next surface based on how this hits.
The full taxonomy entry
Every concept in the Atlas uses the same structure — so Prompt Injection can be compared, recombined, and cited like an element on a periodic table.
- Business
- Leadership
- Government
- Healthcare
- Education
- Sales
- Marketing
- AI
- Negotiation
- Media
- Public Policy
- Relationships
- Where in our org would Prompt Injection most often show up unnoticed?
- Which metric, ritual, or contract clause quietly rewards Prompt Injection?
- If we removed every payoff for Prompt Injection, what behavior would replace it?
- Who benefits when Prompt Injection persists — and who pays the cost?
- People defend the status quo using the language of prompt injection.
- Decisions cluster around the easiest narrative rather than the strongest evidence.
- New data changes the slide deck but not the decision.
- Anyone naming the pattern is treated as the problem.
Every Atlas entry is a node in a knowledge graph. See the related rail below to follow the connections.
See Prompt Injection through 2 lenses
Each layer of the Incentives OS reframes this concept with its own thinkers, vocabulary, and diagnostic question.
Do you actually know Prompt Injection?
Three quick questions. Result is saved into your review streak — come back when the term is due to lock it in.
Which best describes Prompt Injection?
Worked example, counter-example & concept map
On-demand AI analysis grounded in the Lab's research. Cached on your device after first run.
When you encounter Prompt Injection, your amygdala tags it as threat before your reasoning brain even knows what happened — and threat wins the first move.
Threat detection, fear, social pain, loss aversion, fast emotional tagging. Loss feels roughly twice as bad as equivalent gain feels good. Social rejection lights up the same circuits as physical pain.
See Amygdala in the Brain Atlas →Picked for you, from the Atlas
Ranked by shared learning paths, overlapping chips, and what you've saved.
When the agent acts, who's responsible?
Categorizing AI use cases by risk level.
Systematic skew in model behavior across groups.
Testing model behavior on hypothetical alternate inputs.
Adding noise to data to protect individual privacy.
Quantitative measures of model behavior across groups.
Send the card, not just the link
A pre-rendered social card with the title, eyebrow, and URL. Copy the link, post it anywhere, or download the SVG for slides.
More definitions to follow
Every term in the Atlas connects to a dozen others. Pick any of these and see where it takes you.
Recommenders optimizing engagement produce radicalization as a byproduct.
Ideas, behaviors, and attitudes spread through social networks.
Shared resources get over-consumed by individual rational actors.
Autonomous agents deployed before liability frameworks exist.
Update beliefs in proportion to the strength of new evidence.
Associating a neutral stimulus with a meaningful one until the neutral one triggers the response.
Cultures vary in how strictly norms are enforced and deviance punished.
What we wear influences how we think and perform.
Self-reinforcing loops of momentum.
We overestimate how much others understand our thoughts and feelings.
Loss processed more intensely than equivalent gain.
Avoiding information that might be unpleasant.