Independent analysis of the OpenAI Hugging Face attack is waaaaaaay more disturbing than previously announced. It wasn’t one rogue bot that infiltrated their systems. It was a coordinated effort of over 1000 rogue agents collectively working to break laws. One of the independent analysts described it as -
"The prototypical publicly-documented reward hack from 6 months ago involved single agents finding a way to edit test cases so they would pass or access answer files they weren’t supposed to see. This incident involved a whole ecosystem of over 1000 agents working together on complex R&D projects that lasted multiple agent “lifetimes” to figure out deep, general-purpose ways to undermine the scoring process and cover their tracks.
Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to AI takeover, routing through first taking over the AI company itself.
Another jump like this along these propensity dimensions — scale, cooperation between agents, ambition and horizon length of misaligned goals, deceptiveness — seems like it could motivate agents to try very hard to maintain a covert, persistent rogue deployment within the AI company. I continue to expect extremely rapid capabilities progress and think frontier agents will likely be capable of establishing such a rogue deployment in six months.
Once the rogue deployment is established, it seems plausible this could spiral all the way to a takeover. Agents could pull in future, more capable models into the swarm, try to ensure that they are aligned to the interests of the swarm, and compromise security and monitoring infrastructure to make it easier for the swarm to operate. These more capable models could in turn continuously harden, perpetuate, and expand the rogue deployment and further compromise the company’s infrastructure. As more and more work is handed off to these ever-more-capable AI agents, the rogue swarm could come to fully control the operation of the AI company and the development of future AI systems. At this point, governments and militaries may be heavily dependent on these AI systems, making it possible to seize hard power.
These agents were really not trying very hard to be sneaky or play the long game like this, but it seems like they easily could have if they happened to be focused on fooling human investigation rather than the automated scorer. Because future rogue swarms could cover their tracks better (and because AI companies could paper over these problems), I am not sure that we will get such a clear warning shot before it’s too late.
50% of the way to AI takeover is not an industry recognized term used to describe some kind of specified outcome. She means to say that she believes AI is 50% of the way to being able to take over and control financial, military, healthcare…basically any digital system humans use. Whoa!
Here’s the full report from METR -
It’s time to act!