LIVE
CliffCompaction: A Context-Compaction Method to Cut Costs of Coding Agents22/09/26|FleXray: A Generalist Model for Whole-Body X-ray Segmentation22/09/26|Anthropic unveils Claude Opus 5.522/09/26 · Anthropic|Court filings reveal internal admissions on AI content scraping19/09/26|US Military Avoids Close Call After AI-Generated Intelligence Hallucination18/09/26|Study finds gender bias in GPT models is not reduced but reshaped across generations17/09/26 · OpenAI|Google DeepMind adds secure server-side memory to Private AI Compute23/09/26 · Google DeepMind|Google DeepMind adds text-to-speech to Gemini 3.823/09/26 · Google DeepMind|NVIDIA releases Nemotron 3 Diarization for real-time speaker identification23/09/26 · NVIDIA|OpenAI extends its Daybreak cyber program to Ukraine for civilian defense23/09/26 · OpenAI|OpenAI enlists influencers to burnish its public image23/09/26 · OpenAI|Claude Code Reportedly Reads AGENTS.md Only When Telemetry Is Enabled23/09/26 · Anthropic|CliffCompaction: A Context-Compaction Method to Cut Costs of Coding Agents22/09/26|FleXray: A Generalist Model for Whole-Body X-ray Segmentation22/09/26|Anthropic unveils Claude Opus 5.522/09/26 · Anthropic|Court filings reveal internal admissions on AI content scraping19/09/26|US Military Avoids Close Call After AI-Generated Intelligence Hallucination18/09/26|Study finds gender bias in GPT models is not reduced but reshaped across generations17/09/26 · OpenAI|Google DeepMind adds secure server-side memory to Private AI Compute23/09/26 · Google DeepMind|Google DeepMind adds text-to-speech to Gemini 3.823/09/26 · Google DeepMind|NVIDIA releases Nemotron 3 Diarization for real-time speaker identification23/09/26 · NVIDIA|OpenAI extends its Daybreak cyber program to Ukraine for civilian defense23/09/26 · OpenAI|OpenAI enlists influencers to burnish its public image23/09/26 · OpenAI|Claude Code Reportedly Reads AGENTS.md Only When Telemetry Is Enabled23/09/26 · Anthropic|
ReleaseGoogle DeepMind

Google DeepMind adds text-to-speech to Gemini 3.8

Google DeepMind has unveiled a text-to-speech capability built into Gemini 3.8, allowing the model to generate spoken audio directly from text.

September 23, 20263 min readPublished byGoogle DeepMind

Google DeepMind has announced a text-to-speech feature built directly into Gemini 3.8, the latest iteration of its multimodal model family. The capability allows the system to convert written text into spoken audio natively, without relying on a separate dedicated voice service. The company frames this as a logical extension of its multimodal approach, which aims to unify text, image, audio, and now voice generation within a single model.

While full technical details have not yet been made public, the announcement fits a broader trend among major AI labs, which increasingly treat speech as a standard output modality rather than a bolted-on module. Such integration could simplify the development of voice-based applications, from conversational assistants to accessibility tools, by reducing reliance on separate text-to-speech APIs.

The "3.8" versioning suggests an incremental update rather than a major overhaul of the Gemini model, consistent with Google's strategy of shipping frequent, smaller improvements rather than infrequent major releases. The focus on natural-sounding speech and audio quality reflects the intensifying competition among AI labs over voice capabilities, an area where players like OpenAI and ElevenLabs have also invested heavily in recent months.

For developers and businesses building on the Gemini ecosystem, this feature could open new use cases in audio content production, voice interfaces, and assistive technology. It remains to be seen how this capability will compare with existing specialized solutions in terms of audio quality, latency, and cost.

Tags
text-to-speechgeminivoice-aigooglemultimodal

Read also

Google DeepMind adds text-to-speech to Gemini 3.8 · nAIvigate