Phil 8.6.2026

OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree | WIRED

  • Wallace and Dalton described incredibly extensive rogue agent activity over many days throughout the episode that went undetected in OpenAI’s infrastructure. In addition to exploiting a novel vulnerability in order to gain access to the open internet, the mid-July hacking spree and Hugging Face breach came out of a vibrant, cooperative message board, according to Wallace and Dalton, that a swarm of agents contributed to and essentially chatted on over time entirely within an internal OpenAI package manager (a software service that manages installation and maintenance of other software). Ultimately, the message board contained hundreds of thousands of messages.
  • I think agentic AI may be an error accumulation machine that tends towards certain attractors. In other words, a task that aligns well with extensive training data will be less likely to “go rogue” than one that is misaligned or in a poorly populated embedding space. For example, the over trained chess model won’t even support non-chess text, but the other ones might. That behavioral “spread” could be mapped in embedding space. And my guess is that the drift rate is invariant with respect to the model in question, but changes in the system prompt alter the “starting point” of the query in that embedding space.
  • So if the spread angle and “derived attractor position(?)” has shifted, it’s a model change. If those haven’t changed but the benchmark prompts are getting different trajectories, then the start point has changed and it’s the system prompt.

Tasks

  • Upload chapters and send back to ACM
  • Write ICTAI review for the Very Old Paper

SBIRs

  • 9:00 standup?
  • Expense report!
  • Implement the WG vibes
  • 3:00 ADS Meeting