works
Richard Ngo Reframing AGI threat models online Richard Ngo argues that the conventional distinction between AI misuse and misalignment lacks utility for both technical and governance purposes: as systems become more agentic, the way one misuses an AI is essentially to tell it to go off and act autonomously, which collapses the boundary between the two categories. He proposes replacing this binary with the concept of \emphmisaligned coalitions—groups of humans and AI systems jointly pursuing illegitimate power acquisition—and suggests this reframing accommodates diverse threat perspectives, from decentralization risks to concerns about power concentration, potentially unifying fragmented safety research communities. – AI-generated abstract.

Reframing AGI threat models

Richard Ngo

FAR.AI, October 24, 2024

Abstract

Richard Ngo argues that the conventional distinction between AI misuse and misalignment lacks utility for both technical and governance purposes: as systems become more agentic, the way one misuses an AI is essentially to tell it to go off and act autonomously, which collapses the boundary between the two categories. He proposes replacing this binary with the concept of \emphmisaligned coalitions—groups of humans and AI systems jointly pursuing illegitimate power acquisition—and suggests this reframing accommodates diverse threat perspectives, from decentralization risks to concerns about power concentration, potentially unifying fragmented safety research communities. – AI-generated abstract.