RLHF
Reinforcement learning from human feedback.
"We taught the model what we wanted by clicking thumbs."
What is RLHF? Reinforcement learning from human feedback. Quality of feedback shapes quality of model.
How modern LLMs were tuned to follow instructions.
Quality of feedback shapes quality of model.
Diverse, calibrated, well-incentivized feedback labor.
Pick what to reward the model for.
Reinforcement learning from human feedback. In the wild: How modern LLMs were tuned to follow instructions.
Pick a lever. There are no neutral ones — every incentive funds a behavior somewhere.
Pick a reaction to RLHF
One tap. We'll point you at the most useful next surface based on how this hits.
The full taxonomy entry
Every concept in the Atlas uses the same structure — so RLHF can be compared, recombined, and cited like an element on a periodic table.
- Business
- Leadership
- Government
- Healthcare
- Education
- Sales
- Marketing
- AI
- Negotiation
- Media
- Public Policy
- Relationships
- Where in our org would RLHF most often show up unnoticed?
- Which metric, ritual, or contract clause quietly rewards RLHF?
- If we removed every payoff for RLHF, what behavior would replace it?
- Who benefits when RLHF persists — and who pays the cost?
- People defend the status quo using the language of rlhf.
- Decisions cluster around the easiest narrative rather than the strongest evidence.
- New data changes the slide deck but not the decision.
- Anyone naming the pattern is treated as the problem.
Every Atlas entry is a node in a knowledge graph. See the related rail below to follow the connections.
See RLHF through 2 lenses
Each layer of the Incentives OS reframes this concept with its own thinkers, vocabulary, and diagnostic question.
Do you actually know RLHF?
Three quick questions. Result is saved into your review streak — come back when the term is due to lock it in.
Which best describes RLHF?
Worked example, counter-example & concept map
On-demand AI analysis grounded in the Lab's research. Cached on your device after first run.
When you encounter RLHF, your prefrontal cortex has to do extra work to override the automatic response — and that override budget is finite.
Executive control, planning, impulse override, working memory, System 2. First thing to go offline under stress, fatigue, or low blood sugar. Why your 4pm decisions are worse than your 9am ones.
See Prefrontal in the Brain Atlas →Picked for you, from the Atlas
Ranked by shared learning paths, overlapping chips, and what you've saved.
Models trained to follow a written set of principles.
Optimizing a proxy of the goal degrades the actual goal.
Outer: the spec matches our intent. Inner: the model actually pursues the spec.
The trained model develops its own internal optimizer.
What the model is actually optimizing.
Maximizing the reward signal in unintended ways.
Send the card, not just the link
A pre-rendered social card with the title, eyebrow, and URL. Copy the link, post it anywhere, or download the SVG for slides.
More definitions to follow
Every term in the Atlas connects to a dozen others. Pick any of these and see where it takes you.
Executives leave well; employees leave thin.
Stop, Take a breath, Observe, Proceed — a micro-intervention to insert a gap between stimulus and response.
Attacking the person rather than the argument.
Replacing a hard question with an easier one without realizing it.
Underpriced externalities keep dirty energy artificially competitive.
Retroactive extensions privilege legacy estates over public-domain enrichment.
Drawing conclusions about individuals from group-level data.
Training models across devices without centralizing data.
More homework signals rigor to parents but often produces burnout, not understanding.
Three-stage cycle: learn by seeing, learn by doing, learn by teaching — looped continuously.
A bias against ideas or products that originated outside the group.
Individual speed gains hide collective quality decline.