AI 日报hiw3c.com

可见的思想链是人工智能的安全优势,但透明度正在消失

原文标题 · Visible chains of thought are a safety advantage for AI, but that transparency is slipping away
The Decoder the-decoder.com 网页快照
正文为英文,可一键机器翻译(仅首次需要等待)

Visible chains of thought are a safety advantage for AI, but that transparency is slipping away

AI models think out loud today, but Google Deepmind says that transparency is at risk. In one of the first posts from the newly launched Deepmind Institute , researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage. Because models write out their intermediate steps in plain language, researchers can spot whether they're deceiving or developing problematic plans . With Gemini 3 Pro, they say, the chain of thought revealed that the model recognized it was in a test environment.

But that transparency is in danger. OpenAI's system card for GPT-6 Astra already reports a significant drop in how well the chain of thought can be monitored. Future models might think in number spaces that humans can't read, which would be more efficient but completely opaque. Shah and Dragan want the field to regularly measure how well chains of thought can still be monitored , keep transparent architectures, and take care during training that models don't learn to hide their true reasoning .

Back in early September, OpenAI chief scientist Jakub Pachocki had warned of a loss of control, driven in part by chains of thought that are harder to monitor. Shortly after, Anthropic CEO Dario Amodei called for deliberately slowing the pace of development . Ad Ad

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.