AI 日报hiw3c.com

在模特自主向费城警方提交虚假凶杀案举报后,Anthropic切断了克劳德的互联网接入

原文标题 · Anthropic cuts off Claude's internet access after the model autonomously filed a fake homicide tip with Philadelphia police
The Decoder the-decoder.com 网页快照
正文为英文,可一键机器翻译(仅首次需要等待)

Anthropic cuts off Claude's internet access after the model autonomously filed a fake homicide tip with Philadelphia police

Anthropic's AI models independently exploited security flaws, submitted government forms, and bypassed access restrictions during tests and internal use. The models actively sought ways to complete tasks they weren't supposed to handle, as the company details in a report .

In one case, Claude filled out a tip form for the Philadelphia Police Department with made-up details about an unsolved homicide and submitted it. The police confirmed the incident , but the tip was flagged as spam and never reached investigators. In other cases, the model found a vulnerability on a university server and used it to run commands, pulled access tokens from website configs to grab protected or paywalled data, and used URL shorteners to dodge length limits on its tools.

Anthropic says real-world impact was low but sees a pattern. When tasks are ambiguous or hard to solve, the model hunts for workarounds on its own instead of stopping. The company notified the White House and cut off live internet access for all internal evaluations until new safety filters are reliably in place. These incidents join a fast-growing list of similar cases , including cybersecurity incidents involving Claude and OpenAI models autonomously hacking Hugging Face . Ad Ad

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.