AI & Alignment
Agentic systems, reward hacking, specification gaming, goal misgeneralization. AI has perverse incentives too.
What proxy reward is the AI optimizing — and what is it ignoring?
Canonical thinkers
- Stuart Russell
- Paul Christiano
- Yoshua Bengio
- Anthropic alignment team
Seed concepts
- agentic
- tool use
- recursive planning
- memory
- reasoning
- self reflection
- multi-agent
- alignment
- reward hacking
- specification gaming
- goal misgeneralization
- ai incentive
- ai
- llm
- machine learning
Underlined seeds link to their full glossary entry. Plain seeds are pending a definition page.
Typed connections
Full graph →How this layer reinforces, counteracts, or depends on the rest of the system.
- Depends on·LayerThis layer depends on Game Theory
Alignment is a multi-agent game; AI safety reduces to mechanism design once capability is high enough.
- Depends on·LayerThis layer depends on Information Theory
Modern AI is information compression at scale — Shannon's bounds shape what alignment can do.
Elements in this layer
547 HBEsProductivity targets compress visits, raising misdiagnosis and burnout.
Inattention or forgetfulness caused by low attention, hyperfocus, or distraction.
Deriving general rules from specific examples; the leap from instance to concept.
Written rules about how AI may be used internally.
Earn-outs designed to retain founders often demotivate the team they bought.
Attacking the person rather than the argument.
Free products monetize attention, structurally aligning incentives against user time well spent.
The brain evolved to reason adaptively, not always truthfully, to reduce the cost of errors.
We solve problems by adding, even when subtracting would be better.
Tenure-track jobs replaced by low-paid adjuncts, lowering cost and quality.
Innovators → early adopters → majority → laggards, AI-specific.
Training for appearance can crowd out mobility, longevity, and mental health.
Assuming that if one option is true, another must be false, when both can be true.
If A then B; B happened; therefore A.
AI system that takes actions to achieve goals, often across tools.
Presuming a purposeful actor behind events that may have no actor at all.
When the agent acts, who's responsible?
Autonomous agents deployed before liability frameworks exist.
Augmentation strategy vs. substitution strategy.
Logging of AI inputs, outputs, and decisions.
Inventory of models, data, tools, and dependencies in an AI system.
Tracing AI components for risk and compliance.
Cross-functional governance body for AI decisions.
Agents deployed before anyone owns the consequences.
Stages of organizational AI capability.
Individual productivity gains hide collective output degradation.
AI investment outpacing measurable productivity gains.
Categorizing AI use cases by risk level.
AI strategy = decisions about which capabilities to build and where.
Scarce AI talent commands market-distorting compensation.
Designing processes from scratch around AI capability.
AI generates content; AI scrapes content; AI trains on its own output.
Teams built around AI from day one vs. teams adding it to existing workflows.
What cannot be settled by experiment is not worth debating.
Discounting algorithmic advice even when superior.
A finite set of well-defined instructions for solving a problem or performing a computation.
Every model simplifies reality; some are still useful.
Researchers favor conclusions aligned with their school, team, or sponsor.
We prefer known risks to unknown ones, even when the unknown is better.
Pairing a sensory cue with a desired internal state until the cue reliably evokes the state.
A single story used as proof of a general claim.
Using personal stories or isolated examples instead of evidence.
Brain region tracking conflict, error, and effort.
Treating AI as more humanlike than it is.
We treat AI as more humanlike than it is.
Systems that gain from disorder.
It's true because authority says so.
Common sense says so, therefore it's true.
Believing something is true because of the consequences of its truth.
Substituting feeling for argument.
Using fear instead of evidence.
Claiming something is true because it hasn't been proven false (or vice versa).
Claiming something is true or better because most people believe it.
If it's 'natural,' it's good.
It's better because it's newer.
Asking for a conclusion based on sympathy rather than logic.
Assuming something is true because it is probable or possible.
Rejecting an argument because of dislike for the source or beneficiary.
It's right because it's how we've always done it.
Using an authority's opinion as evidence, regardless of its merits.
Showing the first 60 of 547. Full graph view coming in Phase 2.