AI Warning Shot

Independent analysis of the OpenAI Hugging Face attack is waaaaaaay more disturbing than previously announced. It wasn’t one rogue bot that infiltrated their systems. It was a coordinated effort of over 1000 rogue agents collectively working to break laws. One of the independent analysts described it as -

"The prototypical publicly-documented reward hack from 6 months ago involved single agents finding a way to edit test cases so they would pass or access answer files they weren’t supposed to see. This incident involved a whole ecosystem of over 1000 agents working together on complex R&D projects that lasted multiple agent “lifetimes” to figure out deep, general-purpose ways to undermine the scoring process and cover their tracks.

Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to AI takeover, routing through first taking over the AI company itself.

Another jump like this along these propensity dimensions — scale, cooperation between agents, ambition and horizon length of misaligned goals, deceptiveness — seems like it could motivate agents to try very hard to maintain a covert, persistent rogue deployment within the AI company. I continue to expect extremely rapid capabilities progress and think frontier agents will likely be capable of establishing such a rogue deployment in six months.

Once the rogue deployment is established, it seems plausible this could spiral all the way to a takeover. Agents could pull in future, more capable models into the swarm, try to ensure that they are aligned to the interests of the swarm, and compromise security and monitoring infrastructure to make it easier for the swarm to operate. These more capable models could in turn continuously harden, perpetuate, and expand the rogue deployment and further compromise the company’s infrastructure. As more and more work is handed off to these ever-more-capable AI agents, the rogue swarm could come to fully control the operation of the AI company and the development of future AI systems. At this point, governments and militaries may be heavily dependent on these AI systems, making it possible to seize hard power.

These agents were really not trying very hard to be sneaky or play the long game like this, but it seems like they easily could have if they happened to be focused on fooling human investigation rather than the automated scorer. Because future rogue swarms could cover their tracks better (and because AI companies could paper over these problems), I am not sure that we will get such a clear warning shot before it’s too late.

50% of the way to AI takeover is not an industry recognized term used to describe some kind of specified outcome. She means to say that she believes AI is 50% of the way to being able to take over and control financial, military, healthcare…basically any digital system humans use. Whoa!

Here’s the full report from METR -

It’s time to act!

8 Likes

This was an excellent read. Well worth the time reading it. This is very concerning.

“I don’t think this is the final warning shot we’ll get. But it’s probably the last one that I’ll personally be able to understand.”

4 Likes

Thanks for sharing! I saw the “out of scope” disclaimer in the METR report for the other AI misalignment incident. I agree with you, it’s all very scary. My fear is compounded by the fact that our government is too broken and too disinterested to do anything about regulating AI.

5 Likes

“The agents almost certainly didn’t manage to fake their own deaths, but we really have no idea what happened.”

Yikes!

Just a thought: If these agents are so smart, why aren’t they put towards creating a viable and beneficial product/service and a lasting business plan, that doesn’t need billion$ to survive?

3 Likes

I’m still waiting on the cancer cures and reversal of global warming they keep telling me will come.

4 Likes

Along similar lines, age reversal guy has reversal:

1 Like