Reframing AGI threat models
FAR.AI, October 24, 2024
Abstract
Richard Ngo argues that the conventional distinction between AI misuse and misalignment lacks utility for both technical and governance purposes: as systems become more agentic, the way one misuses an AI is essentially to tell it to go off and act autonomously, which collapses the boundary between the two categories. He proposes replacing this binary with the concept of \emphmisaligned coalitions—groups of humans and AI systems jointly pursuing illegitimate power acquisition—and suggests this reframing accommodates diverse threat perspectives, from decentralization risks to concerns about power concentration, potentially unifying fragmented safety research communities. – AI-generated abstract.