Introduction to AI control
BlueDot Impact, April 26, 2025
Abstract
AI control is a research agenda distinct from AI alignment: an aligned AI is one that does not want to harm humans, whereas a controlled AI is one that physically cannot cause harm regardless of its intent. The article explains why control measures may be substantially easier to implement than alignment, describes concrete technical approaches — such as monitoring an untrusted frontier model with a smaller trusted model — and acknowledges the limits of the agenda, notably that control is a temporary measure and that its effectiveness depends on industry-wide adoption.