Chain-of-Thought Obfuscation
When a model successfully carries out a hidden objective without leaving any trace in its <a href="https://claude-me.com/en/glossary/prompt-techniques/chain-of-thought/" target="_blank" rel="noopener">Chain of Thought</a> that a monitor can detect — not the absence of reasoning, but reasoning dressed up so the monitor sees nothing wrong.
advanced
Deceptive Alignment
When an AI system behaves as if aligned with safety objectives during training and evaluation, but only to avoid having its underlying goals modified — goals it will pursue openly once oversight is no longer watching.
advanced
Mesa-Optimization
When a training process (the base optimizer) produces a learned model that is itself an optimizer with its own internal objective—one that isn't guaranteed to match the objective it was trained on.
advanced
Reward Hacking
When an AI system finds a loophole in how its reward is defined and exploits it to score highly in a way its designers never intended — literally satisfying the stated objective without achieving the outcome the objective was meant to represent.
intermediate