Scientific Definition
Cache common prompt prefixes to reduce cost and latency.
Plain-English Definition
Cache common prompt prefixes to reduce cost and latency.
Feynman Explanation
Don't pay twice for the same setup.
Core Principle
Cache common prompt prefixes to reduce cost and latency.
Mechanisms
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Pending editorial review.
Inputs (Triggers)
Pending editorial review.
Outputs (Behaviors)
Pending editorial review.
Behavioral Signature
Don't pay twice for the same setup.
Examples
- Common system prompts cached across calls.
- Cost optimization in production deployments.
Pending editorial review.
Famous Experiments
Pending editorial review.
Design Principles
- Architect for cache reuse.
Measurement Approaches
Pending editorial review.
Evidence
Pending editorial review.
Pending editorial review.
The Perverse Incentive Lens™
How this behavior is exploited — and how to redesign around it.
- Architect for cache reuse.
Pending editorial review.
Pending editorial review.
Interactive Mini Network
Click any neighbor to re-center the graph and follow the threads of connection.
Knowledge Graph Neighbors
Auto-linked to the rest of the Human Behavior Taxonomy by family, domain, dimension, and shared keywords.
AI investment outpacing measurable productivity gains.
GPU access and pricing shape what's feasible.
Operating economics shaped by per-token pricing.
A handful of providers shape the entire AI economy.
Training is a one-time cost. Inference is forever.
Real cost includes data, ops, monitoring, governance, training.
Malicious instructions hidden in user input or retrieved content.
Written rules about how AI may be used internally.
Innovators → early adopters → majority → laggards, AI-specific.
AI system that takes actions to achieve goals, often across tools.
When the agent acts, who's responsible?
Augmentation strategy vs. substitution strategy.