AI 日报hiw3c.com

双子座现场音频

原文标题 · Gemini Live audio
Simon Willison simonwillison.net 网页快照
正文为英文,可一键机器翻译(仅首次需要等待)

Simon Willison’s Weblog

Google released Gemini 3.8 Live and 3.8 Live Extended Thinking today - two new speech-to-speech models that are a similar shape to OpenAI's GPT-Live family.

I pointed GPT-6 Astra Extra High at the documentation and had it build me this web UI for trying out the new models. You can select a model and voice preset, enter an optional system prompt and then start a voice conversation through your browser, including the ability to interrupt the model while it is talking.

The implementation uses no libraries. It connects to the wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=... WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback.

Here's the Gemini Live tutorial for getting started with that WebSockets API.

Recent articles

Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026

OpenAI agents attacked RubyGems back in May - 12th September 2026

Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026

This is a beat by Simon Willison, posted on 15th September 2026 .

Monthly briefing

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.