AI 日报hiw3c.com

双子座3.8 TTC游乐场

原文标题 · Gemini 3.8 TTS Playground
Simon Willison simonwillison.net 网页快照
正文为英文,可一键机器翻译(仅首次需要等待)

Simon Willison’s Weblog

Google released two new Gemini text-to-speech models today - gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts .

They come with a library of over 2,000 voices, plus the ability to create a custom voice with "just a 30-second audio sample of your voice or a voice you have the rights to use".

I vibe coded this bring-your-own-key playground interface with GPT-6 Astra, taking advantage of the open CORS policy of the underlying Gemini API.

A notable feature of the API is that it makes it easy to define a full conversation between multiple characters, each with different voices and voice style instructions.

Here's a short demo clip of a conversation between two pelicans debating if they should move to the Pacifica Pier . I had Claude 4.5 Opus write the script and generate a URL to render it using the tool .

Your browser does not support the audio element.

It took ~20 seconds to generate 1m 18s of audio using Gemini 3.8 Flash TTS (not the cheaper Flash-Lite), at a cost of 2.74 cents.

Recent articles

Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war - 22nd September 2026

Jev introduces a new shape of LLM - System One, aka Decision Models - 21st September 2026

Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026

This is a beat by Simon Willison, posted on 23rd September 2026 .

Monthly briefing

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.