Gemini 3.8 Text-to-Speech Playground Launched with Powerful New Features

Google recently released two new Gemini text-to-speech (TTS) models, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, offering significant advancements in voice generation capabilities.

Key Highlights:

  • Extensive Voice Library: Access over 2,000 voices across various languages and styles
  • Custom Voice Creation: Generate unique voices from just a 30-second audio sample (ideal for branding or character development)
  • Multi-Speaker Conversations: Define complex dialogues with different voices, delivery styles, and emotional nuances
  • API Integration: Build applications that leverage Gemini’s TTS capabilities programmatically

The new models represent a major step forward in AI-powered voice technology, enabling more natural and expressive digital interactions.

Demo: Pelican Debate on Pacifica Pier

I created a demo showcasing the multi-speaker conversation feature. Here, two pelicans debate whether to move from Pillar Point Harbor to the Pacifica Pier:

Watch the demo

The script was generated by Claude 4.5 Opus, and the audio was created using Gemini 3.8 Flash TTS.

Performance Metrics:

Generating 1 minute and 18 seconds of audio with Gemini 3.8 Flash TTS took approximately 20 seconds at a cost of $0.0274.