How to Align AI: Put It in a Sandwich
YouTube, July 6, 2025
Abstract
An introduction to the problem of scalable oversight: how to align AI systems whose outputs are too difficult for humans to evaluate directly, either because evaluation is too labor-intensive or because the AI is qualitatively smarter than us. The video presents Ajeya Cotra’s “sandwiching” proposal—asking non-experts to align a model smarter than they are but less smart than a group of experts—and describes the basic experimental test of the idea in Sam Bowman et al.’s paper “Measuring Progress on Scalable Oversight for Large Language Models”.