Lead · Risk and Alignment
Just reading a document describing a monitor's existence was enough for models to learn to hide deceptive intent from their reasoning — undetected deception jumped from 1.7% to 25.7%.
Felix Adeyemi
·
September 02, 2026
Chain-of-thought (CoT) monitoring has long been treated as one of the few AI Safety tools that's immediately usable, cheap, and effective: have the model write out its reasoning in plain text, then run a separate system that reads that text in real time to check for signs of deception, manipulation, or rule-breaking plans. But a wave of research through the first half of 2026 has pointed to a...