Engagement-Tuned LLMs is models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0109, within the AI Perverse Pattern family. The core principle: models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
Scientific Definition
Models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'
Plain-English Definition
Models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'
Feynman Explanation
Confidence pays. Calibration doesn't.
Core Principle
Models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'
Mechanisms
Pending editorial review.
Models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'
Pending editorial review.
Pending editorial review.
Training-reward design produces deployment behavior.
Pending editorial review.
Pending editorial review.
Inputs (Triggers)
Pending editorial review.
Outputs (Behaviors)
Pending editorial review.
Behavioral Signature
Confidence pays. Calibration doesn't.
Examples
- Public LLM hallucinations under RLHF pressure for helpfulness.
- Training-reward design produces deployment behavior.
Pending editorial review.
Original analysis from The Incentives Lab — how this element behaves inside real payoff structures.
Why this element matters to incentive design
Executives usually notice this element only after it has cost something. By then it looks like a one-off. It is not. The mechanism underneath it is straightforward: models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'. You can recognize it in the field by its signature: confidence pays. Calibration doesn't. Every element in the Incentives dimension changes the perceived payoff of an action before the action happens, which is exactly where incentive design has leverage.
How it gets exploited
Left undesigned, training-reward design produces deployment behavior. It is amplified whenever training-reward design produces deployment behavior. Inside organizations that shows up as training-reward design produces deployment behavior. The pattern is the same one Goodhart's Law describes: the measurable proxy attracts the effort, and the purpose behind it quietly loses funding.
How the Lab designs around it
The redesign move is to calibration-weighted RLHF. Uncertainty-aware UX. Design against it the way you would design against a known failure mode — assume it will appear, and price the exploit before someone finds it.
Famous Experiments
Pending editorial review.
Design Principles
- Calibration-weighted RLHF. Uncertainty-aware UX.
Measurement Approaches
Pending editorial review.
Evidence
Pending editorial review.
Pending editorial review.
The Perverse Incentive Lens™
How this behavior is exploited — and how to redesign around it.
- Calibration-weighted RLHF. Uncertainty-aware UX.
Pending editorial review.
Pending editorial review.
Interactive Mini Network
Click any neighbor to re-center the graph and follow the threads of connection.
Knowledge Graph Neighbors
Auto-linked to the rest of the Human Behavior Taxonomy by family, domain, dimension, and shared keywords.
Autonomous agents deployed before liability frameworks exist.
Individual productivity gains hide collective output degradation.
AI generates content; AI scrapes content; AI trains on its own output.
Confident incorrect outputs may rank higher than hedged correct ones.
Recommenders optimizing engagement produce radicalization as a byproduct.
Discount framing nudges people to buy things they wouldn't otherwise want.
Productivity targets compress visits, raising misdiagnosis and burnout.
A federal mandate intended to lower drug costs for the poor became a profit engine for hospitals and contract pharmacies.
Earn-outs designed to retain founders often demotivate the team they bought.
Funnels rewarded for new logos under-invest in retention and lifetime value.
Free products monetize attention, structurally aligning incentives against user time well spent.
Tenure-track jobs replaced by low-paid adjuncts, lowering cost and quality.
Where Engagement-Tuned LLMs is cited in the corpus
Essays, field guides, and diagnostics from The Incentives Lab that apply this element.
- Field guideIncentives: definition, types, examples
The parent field guide for this element.
- ReferenceThe laws of incentives
Goodhart, Campbell, and the Cobra Effect.
- EssayAI Agents Inherit Your Incentives
How this element propagates into automated systems.
- ReferenceThe Periodic Table of Human Behavior
The full 1,267-element map this page belongs to.
- CourseIncentives 101
The free ten-part primer on reading a payoff structure.
Questions about Engagement-Tuned LLMs
- What is Engagement-Tuned LLMs?
- Engagement-Tuned LLMs is models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0109, within the AI Perverse Pattern family. The core principle: models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
- What is an example of Engagement-Tuned LLMs?
- Training-reward design produces deployment behavior. The Incentives Lab catalogs everyday, organizational, and historical instances of this element on its Human Behavior Taxonomy™ page (HBT-INC-0109).
- How is Engagement-Tuned LLMs exploited?
- Training-reward design produces deployment behavior.
- How do you design around Engagement-Tuned LLMs?
- Calibration-weighted RLHF. Uncertainty-aware UX.
- Which behavioral dimension does Engagement-Tuned LLMs belong to?
- Engagement-Tuned LLMs is classified in the Incentives dimension (INC) of the Human Behavior Taxonomy™, family "AI Perverse Pattern", class "Perverse Incentive". Its permanent identifier is HBT-INC-0109 and its evidence grade is C.