works
Sella Nevo et al. A playbook for securing AI model weights online This RAND research brief summarizes recommendations for securing the weights of frontier AI models against theft and misuse. Because stealing model weights lets attackers exploit a model for their own ends, the authors define five security levels ranging from defense against opportunistic attackers up to the most capable nation-state adversaries and map each level to concrete measures — including hardened software interfaces, drastically reducing the set of individuals with copy access, monitoring the full software and hardware supply chain, and using trusted execution environments on GPUs. It offers developers and policymakers a way to match protection to the risk posed by a given system.

A playbook for securing AI model weights

Sella Nevo et al.

RAND Corporation, November 21, 2024

Abstract

This RAND research brief summarizes recommendations for securing the weights of frontier AI models against theft and misuse. Because stealing model weights lets attackers exploit a model for their own ends, the authors define five security levels ranging from defense against opportunistic attackers up to the most capable nation-state adversaries and map each level to concrete measures — including hardened software interfaces, drastically reducing the set of individuals with copy access, monitoring the full software and hardware supply chain, and using trusted execution environments on GPUs. It offers developers and policymakers a way to match protection to the risk posed by a given system.