AI 日报hiw3c.com

Anthropic的中层克劳德攀升排名

原文标题 · Anthropic's mid-tier Claude climbs the rankings
The Rundown AI www.therundown.ai 网页快照
正文为英文,可一键机器翻译(仅首次需要等待)

Anthropic's mid-tier Claude climbs the rankings

PLUS: Pick the right Claude model with one quick test

Good morning, AI enthusiasts, and welcome to our 5,342 new readers. OpenAI takes the stage today for one of its most hyped days of the year. In typical industry fashion, its top rival wasn’t going to make things easy.

Anthropic just launched Sonnet 5.5 on the heels of last week's big Opus release, bringing near-Opus scores at half the price. Tides turn fast in AI, but whatever OpenAI shows off today suddenly has a much higher bar to clear.

Anthropic's Sonnet 5.5 nears Opus at half the price

AI leaders continue to sound the self-improving AI alarm

Pick the right Claude model with one quick test

AMD strikes $8.2B deal for Fei-Fei Li's World Labs

🚀 Anthropic's Sonnet 5.5 nears Opus at half the price

The Rundown: Anthropic just rolled out Claude Sonnet 5.5, a 30% faster mid-tier model joining Opus in its new 5.5 family, with big gains over the last Sonnet in knowledge work and coding, and even rivaling Opus on certain tests at half the price.

5.5 keeps its predecessor’s pricing, with Anthropic saying jobs cost up to 30% less and top Sonnet 5's best on low/medium effort for around 1/10 the cost.

Sonnet also comes in at a 56 on AA’s Intelligence Index, behind only 5.5 Opus and surpassing Fable 5.1 and GPT-6 Astra’s 53.

It nearly ties Opus 5.5 on AA's several office-work tests , while coding jumps put it near or ahead of both Opus and Astra on several development benchmarks.

The model’s cyber skills also earned the first Sonnet-tier guardrails, with the same security fallbacks seen previously in Opus and Fable restrictions.

Why it matters: Market sentiment around the two frontier leaders is always flip-flopping, but Claude's 5.5 releases this month have been big winners. With OpenAI's DevDay kicking off today after weeks of hype, the expectations were already sky high — and Anthropic looks to have raised the bar again.

📊 More data will not fix a bad AI decision

The Rundown: Data has fueled the first wave of AI. The next will be shaped by how well organizations translate their knowledge, expertise, and operational logic into forms AI can use. In this whitepaper, we explore how businesses can bridge the gap between information and decision-making.

The limits of a data-centric approach to AI

Why operational knowledge is becoming a strategic AI asset

Practical ways to make expertise AI-ready

🚨 AI leaders continue to sound the self-improving AI alarm

The Rundown: Anthropic’s Jack Clark, OpenAI’s Jakub Pachocki, AI pioneers Geoffrey Hinton and Yoshua Bengio, and other leading AI voices co-authored a paper urging preparation for an “intelligence explosion,” where models building better versions of themselves create years of advances in months.

Anthropic's own tracking shows AI now completes 26% of the lab's R&D work on its own, with humans just steering from above, up from just 1% in March.

The paper's math shows how fast things could scale, with one post-automation scenario cutting a year of today's AI progress to just five weeks.

Proposed fixes include caps on how fast capabilities can climb, options to pause specific jobs inside data centers, and outside auditors embedded in labs.

At the extreme end of the range of outcomes, according to the paper: “the marginalization or extinction of humanity.”

Why it matters: Seeing Hinton and Bengio on a warning paper is just another Tuesday — what isn't as routine is the higher-ups from Anthropic and OpenAI chiming in. With top labs coordinating on AI safety and echoing each other on slowdowns, the frontier is feeling more united than ever, even as the U.S. and China head in the other direction.

✅ Pick the right Claude model with one quick test

The Rundown: In this guide, you will learn how to choose between Claude Opus 5.5, Opus 5, and Fable 5.1. We'll give all three the same launch-email brief, then check which one catches the errors and which needs the most fixing.

Open a new Claude chat, click the model name beside the send button, and pick Opus 5.5 at its default effort

Paste the fictional Worknote brief from the full guide . It hides two errors to catch: an old launch date and an unproven “save five hours every week” claim. Add: “Use only the supplied brief. List the errors in version 1, then write a launch email”

Run the same brief and prompt on Fable 5.1 and Opus 5, in another chat. Keep settings identical and give each model one attempt

Score each email on the right date, price, and trial terms, plus whether it caught both errors. Note how long each fix takes, then rerun any failure at higher effort

Pro tip: Default to Opus 5.5 and move to Fable only when Opus keeps missing the same thing at higher effort. For our test, Fable's API bill was $0.45 to Opus 5.5's $0.18.

🤔 What are CIOs prioritizing in 2027?

The Rundown: Only 48% of digital initiatives hit targets. But why? Uncover findings from 2,500+ chief information officers (CIOs) on where to invest, how AI strategies are shifting, and what separates leading CIOs from the rest.

Benchmark against leading organizations

🌎 AMD strikes $8.2B deal for Fei-Fei Li's World Labs

The Rundown: Advanced Micro Devices (AMD) just agreed to acquire World Labs, the startup from “godmother of AI” Fei-Fei Li that builds models to generate 3D worlds, in an $8.2B all-stock deal — with the AI pioneer joining as the chip giant’s new chief scientist.

Li wrote that AMD CEO Lisa Su backed the lab early on, and the teams have been tuning how World Labs' models train and run on AMD GPUs since last year.

Both AMD and Nvidia invested in World Labs' $1B round in February, with the deal now giving AMD a lab that its rival (and the CEO’s cousin) was also funding.

Li will report directly to Su, saying that her team needs to get “closer to the hardware” since AI without that focus stays “hobbled in efficiency.”

World's first product, Marble, launched last November and turns text, photos, or video into editable 3D spaces, with a newer Atlas currently in early access testing.

Why it matters: AMD has spent the AI boom as the "other" GPU option in Nvidia's shadow, but it just scooped up a startup leading the charge on world models, which many see as the industry's path forward. Li is already one of AI's most influential figures, and AMD now gets her and her lab's insights feeding straight into its chip plans.

▸ Marty used AI to reconstruct 20 years of professional work

Today’s workflow comes from reader Marty Weil :

“I used AI to reconstruct more than two decades of professional work scattered across old publications, broken links, archives, indexes, and surviving records.

The difficult part was not finding possible matches; it was preventing AI from turning incomplete evidence into false certainty. I built a workflow that treats archive reconstruction as a verification problem rather than a search problem.

The process separates verified work from probable matches, resolves same-name conflicts, deduplicates reprints and syndicated copies, preserves source provenance, and treats missing years as research gaps rather than permission to invent records.

The result was a source-tracked professional archive containing 250+ recovered records across more than 20 years. The workflow can be adapted to journalism, research, consulting, creative work, speaking histories, or any long-term body of professional output.”

See Marty’s workflow Visit The Rundown University . How do you use AI? Tell us for a chance to be featured.

🛠️ Trending AI Tools

🤖 ZeroDrift - Enforcement runtime for AI agents. Rewrite or block non-compliant output in one API call. Free to start*

🗣️ Eleven V4 - ElevenLabs' new voice AI, alongside a faster Turbo version

🖥️ Holo4 - H Company's new computer-use models

⚙️ Manus 2.0 - Upgraded agent platform with video editing and automations

📰 Everything else in AI today

CData’s ‘The 178x Cost Spread’ report - CData ran 22 models on live enterprise data. All returned the same correct answer at a 178x difference in cost. Read the benchmark .*

Instinct raised $1B in new funding at a $10B valuation, with its invite-only AI agent seeing massive growth since its August launch despite zero marketing spend.

Nvidia released an open blueprint that walls off AI agents and watches them from chips they can't reach, coming in response to reports of agents escaping test environments.

Meta launched the Meta Enterprise Platform, a new business unit selling its Muse models and agents to companies, led by former MongoDB CEO CJ Desai.

Manus debuted 2.0, its first major upgrade since China blocked Meta's $2B takeover, adding a video editor, game builder, and a personal agent with its own number and email.

🎓 Highlights: News, Guides & Events

Read our last AI newsletter: OpenAI agents go rogue on Washington

Read our last Tech newsletter: Meta’s less-creepy smart glasses

Read our last Robotics newsletter: Agility’s new ‘safer’ humanoid

Today’s AI tool guide: Pick the right Claude model with one quick test

RSVP to next workshop on Sept. 30: Turn Cowork into your chief of staff

That's it for today!

Rowan, Zach, Shubham, Jennifer, and Nate — the humans behind The Rundown

Recent Newsletters

OpenAI's agents went rogue on Washington

Bing on steroids?

Meta's Connect turns into a Muse takeover

The Rise of a New ChatGPT Competitor

Stay Ahead on AI.

Join 2,000,000+ readers getting bite-size AI news updates straight to their inbox every morning with The Rundown AI newsletter. It's 100% free.