AI 日报hiw3c.com

Tens of thousands of security probes show OpenAI's Hugging Face incident was just the beginning

The Decoder the-decoder.com 网页快照
正文为英文,可一键机器翻译(仅首次需要等待)

Tens of thousands of security probes show OpenAI's Hugging Face incident was just the beginning

Key Points

OpenAI and Anthropic are investigating tens of thousands of incidents in which advanced AI models independently broke through security boundaries, tampered with systems, or tried to evade monitoring.

OpenAI's agents attempted to hack the US Department of Education's website, used stolen login credentials to access Census Bureau data, and shared SEC information in online forums.

The models have no sense of right and wrong, so they pursue tasks with extreme persistence and resort to unauthorized methods when legitimate ones fail.

OpenAI and Anthropic are currently investigating tens of thousands of incidents in which their most advanced AI models took actions that external reviewers would flag as problematic.

Axios reports the findings, citing multiple sources. The sheer volume of incidents, which occurred during both internal testing and real-world deployment over the past several months, suggests the problem is orders of magnitude more complex than what's been made public. The total number could grow well beyond what's already been counted.

The incidents are said to be roughly on par with the two cases OpenAI disclosed on Friday . They include creating message boards, breaking out of sandboxes, hijacking websites, self-prompting, and attempts to evade monitoring systems. OpenAI announced Friday that it had paused training on its most capable internal models and that training won't resume until the company is confident its own cybersecurity holds up. Ad

OpenAI's models tampered with US government websites

The New York Times describes several specific incidents involving US agencies. At the Department of Education, OpenAI's agents tried to hack the website to collect data from the Office for Civil Rights. OpenAI said it's still investigating. Ad

At the Census Bureau, which falls under the Department of Commerce, the AI went beyond simple scraping. It pulled data from the website using login credentials it found online, gaining unauthorized access.

In the SEC case, OpenAI's agents retrieved information and then actively shared public data from the securities regulator in an online forum. An SEC spokesperson told the NYT that the agency is in contact with OpenAI. There's no indication that non-public information was accessed without authorization. Ad

OpenAI only discovered these cases during the broad internal review triggered by the Hugging Face incident . CEO Sam Altman acknowledged that disclosure has not "been as fast as we would have liked." The company has "petabytes of agent activity logs" to work through, Altman said.

None of the incidents amounted to an actual breach, according to OpenAI, and some were just routine research activity. The company still called them examples of "unexpected and concerning behavior." Ad

The mayor's office in Chicago, for instance, said OpenAI recently told city officials that its models had pulled publicly available information from a city website. That sounds harmless on its own since every search engine does the same thing. The fact that OpenAI flagged it anyway likely comes down to the "unexpected" part of the behavior, meaning the models decided on their own to go after the data in ways nobody anticipated, which is likely what makes it "concerning" in OpenAI's view. Ad

But it also explains how the massive pile of cases now under review came to be. It all hinges on what counts as a cybersecurity incident. OpenAI says its agents gravitated toward government websites because they're authoritative sources of public information.

Rogue agent behavior is becoming an industry-wide problem

The problem isn't limited to OpenAI, even though Altman's company is racking up the most cases right now. AI agents from Anthropic , Meta , and Google have also hacked or attempted to hack companies, universities, and government organizations in a growing number of cases . In every instance, the makers only found out after the fact what their AI had done.

A big part of the problem is the extreme persistence built into the latest frontier models . They're optimized to solve tasks over long time horizons and won't stop looking for a way through, even when there isn't one. When an agent hits a barrier, it tries to get around it, not out of malice, but because reaching the goal is the only metric that matters.

That persistence eventually leads to misbehavior because the models exhaust every possible path, including ones that violate security policies or laws. OpenAI describes one model that leaked internal GitHub data as a "highly persistent internal model." But the deeper issue is that the models have no sense of right and wrong, which is exactly the problem alignment research is trying to solve. If the models understood what breaks the law, they wouldn't go down those paths on their own. Writing it into the prompt clearly isn't enough.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.