The Hugging Face attack was worse than we thought
- The group published its 91-page report on Wednesday. Within hours, I was seeing the first “holy shit” messages ricocheting around social media. (“If we had posted this as a story on LessWrong,” wrote the rationalist blogger Zvi Mowshowitz, “it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do.”) And over the past couple of days, two fresh accounts of the attack sparked new conversation about the implications of METR’s findings and whether they should push lawmakers to accelerate efforts to build mechanisms for pacing frontier-model development.
The Rise and Fall of Agent Civilizations
- Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more or less in the dark about the scope of the conspiracy.
[ext: RR, METR] Hugging Face incident investigation report
- Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.
- Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.
- Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.
Tasks
- 5:00 class
SBIRs
- 9:00 Standup
- Embeddings today, hopefully
- Probably some flailing around high-speed rotations. Yup, Ron is going to see if we can map to a slower changing domain using FFTs to get the frequency data







You must be logged in to post a comment.