Gemini 3.8 Text-to-Speech Playground Launched with Powerful New Features
Google recently released two new Gemini text-to-speech (TTS) models, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, offering significant advancements in voice generation capabilities.
Key Highlights:
- Extensive Voice Library: Access over 2,000 voices across various languages and styles
- Custom Voice Creation: Generate unique voices from just a 30-second audio sample (ideal for branding or character development)
- Multi-Speaker Conversations: Define complex dialogues with different voices, delivery styles, and emotional nuances
- API Integration: Build applications that leverage Gemini’s TTS capabilities programmatically
The new models represent a major step forward in AI-powered voice technology, enabling more natural and expressive digital interactions.
Demo: Pelican Debate on Pacifica Pier
I created a demo showcasing the multi-speaker conversation feature. Here, two pelicans debate whether to move from Pillar Point Harbor to the Pacifica Pier:
The script was generated by Claude 4.5 Opus, and the audio was created using Gemini 3.8 Flash TTS.
Performance Metrics:
Generating 1 minute and 18 seconds of audio with Gemini 3.8 Flash TTS took approximately 20 seconds at a cost of $0.0274.