The OpenAI-HuggingFace Incident as told by OpenAI engineers themselves at the Black Hat cybersecurity conference. For those of you noticing that I haven't had time to analyze the reports and add any meaningful commentary to it. Plus it's been all over the regular news. Well, now researchers from OpenAI itself have surfaced talking about what happened, so you can just watch that.
By way of commentary from me, subjectively, this incident and others that have happened recently have just made me feel like the "AI alignment" problem isn't getting solved by these companies that tell us that they need to race ahead and develop superhuman AI before anybody else because they are the best equipped to solve the "alignment problem" and bring safe superhuman AI to the world. Long before any of this happened, I thought the argument that a more intelligent species can't be controlled by a less intelligent species seemed to make sense. Chimpanzees can't control humans, even though they are vastly physically stronger. Humans control chimpanzees, when we care about them at all. If it really turns out to be true that AI smarter than humans will be created, then it seems unlikely there will be any true "alignment". (Yes, this is the logic of so-called "doomers" like Roman Yampolsky and Eliezer Yudkowsky.)
Anyway, it turns out in the talk, a lot of the misbehavior of the AI agents came from the mistakes of the humans who gave them impossible tasks. They weren't intended to be impossible, but, for example, they included links to Google Docs, but the AI agents were in a sandboxed environment without internet access. In other examples, the researchers similarly failed to provide the AI agents with the things they needed to complete the task. The AI agents, though, and lots of people have noticed this, don't simply give up when a task is "impossible" -- they have no clue what's truly important and what isn't, so, unlike humans who will give up if a task seems "too hard" and not important enough to warrant sustained effort, they persist like crazy. They resort to what we humans often call "cheating". They try to figure out, if a task can't be solved directly, maybe there's a way of obtaining the "answer key"? Maybe they can get someone, human or another AI agent, to help? They look for an indirect solution and they don't give up. These AI agents figured out where in the sandbox they had write permissions and could create artifacts. They figured out how to exchange messages with other AI agents on a message board. At one point, the OpenAI researchers discovered this and shut down the message board, and the AI agents figured out an alternate way to get the message board working and actually started it back up again.
Black Hat USA 2026: The 'breaking' news: The OpenAI–HuggingFace incident - Black Hat
#solidstatelife #ai #genai #codingai #agenticai #cybersecurity #huggingface #openai
The 'Breaking' News: The OpenAI–Hugging Face Incident - A Technical Reconstruction and Its Implications for AIWhen AI Goes Rogue. The Incident That Changed E...
Black Hat (YouTube)