AI Alignment & Incentives Catalog
85 terms across 9 themes — every definition in the Atlas sourced from this manual, with executive examples and counter-moves.
Written rules about how AI may be used internally.
Innovators → early adopters → majority → laggards, AI-specific.
AI system that takes actions to achieve goals, often across tools.
When the agent acts, who's responsible?
Augmentation strategy vs. substitution strategy.
Logging of AI inputs, outputs, and decisions.
Inventory of models, data, tools, and dependencies in an AI system.
Tracing AI components for risk and compliance.
Cross-functional governance body for AI decisions.
Agents deployed before anyone owns the consequences.
Stages of organizational AI capability.
AI investment outpacing measurable productivity gains.
Categorizing AI use cases by risk level.
AI strategy = decisions about which capabilities to build and where.
Scarce AI talent commands market-distorting compensation.
Designing processes from scratch around AI capability.
Teams built around AI from day one vs. teams adding it to existing workflows.
Discounting algorithmic advice even when superior.
Treating AI as more humanlike than it is.
Augment when judgment matters. Automate when scale matters.
Reduced vigilance with automated systems.
Systematic skew in model behavior across groups.
Strategic choice on AI capability sourcing.
AI capability outpacing organizational ability to use it.
AI bolted onto existing workflows to look forward-leaning.
GPU access and pricing shape what's feasible.
The relationship between inputs and outputs changes.
Quantifying uncertainty in model outputs.
Models trained to follow a written set of principles.
Cryptographic tracking of content origin.
How much input the model can process at once.
Operating economics shaped by per-token pricing.
Testing model behavior on hypothetical alternate inputs.
Underlying data distribution changes over time.
Tracking where training and inference data came from.
Adding noise to data to protect individual privacy.
Vector representation of content for similarity and search.
Comprehensive AI regulation in the EU.
Evaluation suites that no longer reflect real-world conditions.
Systematic testing of model quality, safety, and capability.
Why did it produce this? vs. How does it work?
Quantitative measures of model behavior across groups.
Training models across devices without centralizing data.
Specialized training on domain data.
A handful of providers shape the entire AI economy.
Optimizing a proxy of the goal degrades the actual goal.
Confident outputs that are factually wrong.
Human review at critical AI decision points.
Human oversight without per-decision review.
Models learning from examples in the prompt.
Training is a one-time cost. Inference is forever.
Outer: the spec matches our intent. Inner: the model actually pursues the spec.
Bypassing model safety constraints.
Redesigning roles around AI capability.
How much delay the user experience tolerates.
The trained model develops its own internal optimizer.
Documentation of model purpose, performance, limitations, and risks.
Train a smaller model to imitate a larger one.
Performance degradation as real-world data shifts.
Multiple specialized agents coordinating on tasks.
U.S. voluntary AI risk management framework.
What the model is actually optimizing.
Where the model runs shapes privacy, latency, cost, and capability.
Open-weight vs. API-only models.
AI pilots that succeed and never scale.
Individual speed gains hide collective quality decline.
Cache common prompt prefixes to reduce cost and latency.
Malicious instructions hidden in user input or retrieved content.
Adversarial testing of AI systems.
Pre-deployment analysis of who loses what.
Investing in workforce transition vs. workforce change.
Generate based on retrieved documents, not just trained weights.
Maximizing the reward signal in unintended ways.
Reinforcement learning from human feedback.
Teams under-reporting AI capability to protect comp or status.
Employees using unauthorized AI tools to get work done.
Foundational skills erode through AI offloading.
The model achieves the goal as stated, not as intended.
AI doesn't just replace tasks — it threatens identities.
Generated training data that mimics real data.
Models calling external functions and APIs.
Real cost includes data, ops, monitoring, governance, training.
Matching trust in a system to its actual reliability.
Database optimized for high-dimensional similarity search.
Concentration risk on a single AI provider.