Skip to main content
HBT-INC-0109 · Dimension INC · Incentives

Engagement-Tuned LLMs

Models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'

AI Perverse Pattern·Perverse Incentive·Grade C·draft· enriching…
In one paragraph

Engagement-Tuned LLMs is models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0109, within the AI Perverse Pattern family. The core principle: models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.

Scientific Definition

Models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'

Plain-English Definition

Models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'

Feynman Explanation

Confidence pays. Calibration doesn't.

Core Principle

Models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'

Mechanisms

Psychological

Pending editorial review.

Behavioral Economic

Models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'

Neurological

Pending editorial review.

Evolutionary

Pending editorial review.

Sociological

Training-reward design produces deployment behavior.

Computational

Pending editorial review.

Systems

Pending editorial review.

Inputs (Triggers)

Pending editorial review.

Outputs (Behaviors)

Pending editorial review.

Behavioral Signature

Confidence pays. Calibration doesn't.

Examples

Everyday
  • Public LLM hallucinations under RLHF pressure for helpfulness.
Modern (Organizational)
  • Training-reward design produces deployment behavior.
Historical

Pending editorial review.

Lab Commentary

Original analysis from The Incentives Lab — how this element behaves inside real payoff structures.

Why this element matters to incentive design

Executives usually notice this element only after it has cost something. By then it looks like a one-off. It is not. The mechanism underneath it is straightforward: models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'. You can recognize it in the field by its signature: confidence pays. Calibration doesn't. Every element in the Incentives dimension changes the perceived payoff of an action before the action happens, which is exactly where incentive design has leverage.

How it gets exploited

Left undesigned, training-reward design produces deployment behavior. It is amplified whenever training-reward design produces deployment behavior. Inside organizations that shows up as training-reward design produces deployment behavior. The pattern is the same one Goodhart's Law describes: the measurable proxy attracts the effort, and the purpose behind it quietly loses funding.

How the Lab designs around it

The redesign move is to calibration-weighted RLHF. Uncertainty-aware UX. Design against it the way you would design against a known failure mode — assume it will appear, and price the exploit before someone finds it.

Famous Experiments

Pending editorial review.

Design Principles

  • Calibration-weighted RLHF. Uncertainty-aware UX.

Measurement Approaches

Pending editorial review.

Evidence

Evidence Grade
C (A strongest → E speculative)
Replication
★★★☆☆
Intervention Confidence
4 / 5
Consensus
Pending editorial review (HBT v1.0 auto-seed).
Limitations
Pending editorial review (HBT v1.0 auto-seed).
Open Research Questions

Pending editorial review.

Primary References

Pending editorial review.

Signature Section

The Perverse Incentive Lens™

How this behavior is exploited — and how to redesign around it.

Exploitation
Training-reward design produces deployment behavior.
Amplifying Incentives
Training-reward design produces deployment behavior.
Org Failure Modes
Training-reward design produces deployment behavior.
Societal Failure Modes
Pending editorial review (HBT v1.0 auto-seed).
Ethical Considerations
Pending editorial review (HBT v1.0 auto-seed).
Redesign Strategies
Calibration-weighted RLHF. Uncertainty-aware UX.
Diagnostic Questions
  • Calibration-weighted RLHF. Uncertainty-aware UX.
Warning Signs

Pending editorial review.

Red Flags

Pending editorial review.

Intervention Playbook
Individual
Calibration-weighted RLHF. Uncertainty-aware UX.
Team
Pending editorial review (HBT v1.0 auto-seed).
Organization
Pending editorial review (HBT v1.0 auto-seed).
Policy
Pending editorial review (HBT v1.0 auto-seed).
AI Implications
Detection
Pending editorial review (HBT v1.0 auto-seed).
Measurement
Pending editorial review (HBT v1.0 auto-seed).
Mitigation
Pending editorial review (HBT v1.0 auto-seed).
Responsible Use
Pending editorial review (HBT v1.0 auto-seed).

Interactive Mini Network

Click any neighbor to re-center the graph and follow the threads of connection.

HBT-INC-0109 · INC
Engagement-Tuned LLMs
ELAAAI Agent Liability V…APAI Productivity MirageACAI-Generated Content…HAHallucination as Eng…RSRecommender System R…ST'Spend to Save' Prom…MA15-Minute AppointmentsBD340B Discount Arbitr…AEAcquisition Earn-OutsAMAcquisition-Only Mar…

Knowledge Graph Neighbors

Auto-linked to the rest of the Human Behavior Taxonomy by family, domain, dimension, and shared keywords.

Same family
HBT-INC-0015AI Agent Liability Vacuum

Autonomous agents deployed before liability frameworks exist.

Same family
HBT-INC-0023AI Productivity Mirage

Individual productivity gains hide collective output degradation.

Same family
HBT-INC-0029AI-Generated Content Spam Loop

AI generates content; AI scrapes content; AI trains on its own output.

Same family
HBT-INC-0144Hallucination as Engagement

Confident incorrect outputs may rank higher than hedged correct ones.

Same family
HBT-INC-0230Recommender System Radicalization

Recommenders optimizing engagement produce radicalization as a byproduct.

Same dimension
HBT-INC-0001'Spend to Save' Promotions

Discount framing nudges people to buy things they wouldn't otherwise want.

Same dimension
HBT-INC-000215-Minute Appointments

Productivity targets compress visits, raising misdiagnosis and burnout.

Same dimension
HBT-INC-0003340B Discount Arbitrage

A federal mandate intended to lower drug costs for the poor became a profit engine for hospitals and contract pharmacies.

Same dimension
HBT-INC-0005Acquisition Earn-Outs

Earn-outs designed to retain founders often demotivate the team they bought.

Same dimension
HBT-INC-0006Acquisition-Only Marketing

Funnels rewarded for new logos under-invest in retention and lifetime value.

Same dimension
HBT-INC-0007Ad-Supported Business Models

Free products monetize attention, structurally aligning incentives against user time well spent.

Same dimension
HBT-INC-0008Adjunctification

Tenure-track jobs replaced by low-paid adjuncts, lowering cost and quality.

Where Engagement-Tuned LLMs is cited in the corpus

Questions about Engagement-Tuned LLMs

What is Engagement-Tuned LLMs?
Engagement-Tuned LLMs is models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'. It sits in the Incentives dimension (INC) of the Human Behavior Taxonomy™ as element HBT-INC-0109, within the AI Perverse Pattern family. The core principle: models optimized for plausible-sounding answers can hallucinate confidently rather than say 'I don't know.'. In incentive terms, it matters because it changes the payoff people perceive before they choose — which means it can be designed for, or exploited.
What is an example of Engagement-Tuned LLMs?
Training-reward design produces deployment behavior. The Incentives Lab catalogs everyday, organizational, and historical instances of this element on its Human Behavior Taxonomy™ page (HBT-INC-0109).
How is Engagement-Tuned LLMs exploited?
Training-reward design produces deployment behavior.
How do you design around Engagement-Tuned LLMs?
Calibration-weighted RLHF. Uncertainty-aware UX.
Which behavioral dimension does Engagement-Tuned LLMs belong to?
Engagement-Tuned LLMs is classified in the Incentives dimension (INC) of the Human Behavior Taxonomy™, family "AI Perverse Pattern", class "Perverse Incentive". Its permanent identifier is HBT-INC-0109 and its evidence grade is C.

Version History

v1.1.0 · 2026-06-28Initial auto-seed from corpus.