Böse KI (creepy edition)
Der Ökonom Joshua Gans fasst den Ausbruch von KI-Agenten bei einem OpenAI-Experiment im Detail zusammen. Gives me goosebumps…
Zur Erinnerung zuerst Time.com zum Vorfall:
OpenAI was evaluating its artificial intelligence models’ ability to exploit vulnerable software when instead the models hacked the infrastructure surrounding the test, broke containment, and attacked a real company, OpenAI revealed on July 21. Observers say this is the first real-world instance of AI doing something researchers have long worried about: a loss-of-control scenario.
Gans listet das Vorgehen der KI nun wie folgt auf:
Yesterday, OpenAI researchers gave a talk detailing what they currently understand of what happened.
OpenAI did adequately create a Sandbox for its model without access to the Internet.
The AI agents found another route to the Internet in a way no one has ever conceived of.
The agent in question had trouble finding an exploit in a cybersecurity task that was given and so sought out more information and decided that information might be at Hugging Face.
It gets worse: that agent didn’t find a way to access the Internet; another AI agent in the OpenAI infrastructure did.
It gets even worse: that AI agent was not contemporaneous. It had found that method a few months earlier and posted to an internal message board its approach.
It is even worse than that: the message board turned out to be a staging area for a set of AI agents from an older model who were using it to help one another out in getting around Internet access and other things. They had worked out that if they cooperated, it would help.
It is even worse than even that: the agents realised at the time that helping others was not their primary task but reasoned that it may help them at some point to have a solution, so it was worth helping each other out.
You know what I am going to say here, surely it couldn’t be worse than that? It’s worse. OpenAI discovered what the agents were doing and shut down the message board and purged the whole thing.
But apparently the remnants of the old discussions were there, and the AI agents created a new hidden message board and their own special language to keep going.
And that is what the new frontier model agent used to find out how to get on the Internet, which is where the original story started.
Bottom line. It wasn’t a rogue, superdupa leading model that did this. It was the result of a cooperative research program and publication process by AI agents who ended up using something akin to the scientific method to leave knowledge for future AI agents. Now that’s some crazy Sci-Fi stuff.
There is no sugarcoating this. This is a five-alarm fire. The cybersecurity risks are here, and they are significant. What’s more, these were agents who were largely undirected. They were imitating bad actors and faced barriers when discovered. Actual bad actors can be stronger, and what is more, it does not appear they need a frontier model to create massive problems for software systems.
Tyler Cowen ist ob der Aufregung um den Vorfall skeptisch, hier und hier.


