Most software days still start the same way. A photo lives in one window. A project board lives in another. A script that needs a voice sits in a third. The work is related, but the tools are not, so the person doing the work becomes the courier between them.
Google moved two pieces of that courier job into Gemini at once.
In the Gemini app, a new set of Connected Apps began rolling out so a request can stay in chat while the action lands in another service.
In Google AI Studio and the Gemini API, two text-to-speech models arrived so a script can stay in the same toolchain and come back as audio.
One change routes language into other software. The other turns language into speech.
They landed on the same day and point at the same shift: Gemini as a place work starts, not only a place questions are answered.
The consumer side is the Connected Apps expansion. Users can link supported services from settings, by using @ mention, or by naming the app in a prompt, then keep going without opening a new tab for every step.
Google grouped the additions into three categories.
Productivity covers Airtable, Linear, monday.com, PandaDoc, Wispr AI, and Zoho. Creativity covers Adobe, Picsart, Squarespace, and Webflow. Lifestyle covers apartments.com, Experian, Peloton, and SeatGeek. The official Gemini account called the same rollout 13 new connections and showed sample prompts: adjust lighting in a photo through Adobe, check what changed in a credit file through Experian, search two-bedroom rentals in Miami through apartments.com, find a 30-minute strength class through Peloton, or create a task on a monday.com board. The extra name in the company blog is Zoho.
The Gemini account also referred to the links as MCP connections. Availability began the same day as the post and was described as a rollout, not an instant global switch.
None of this removes the need to review what the model did in the connected app. It just changes where the request is made.
The developer side arrived a few hours earlier from Google AI Studio: 'Gemini 3.8 Flash TTS' and 'Gemini 3.8 Flash-Lite TTS.'
In a blog post, Google says that the models take text and return audio. They are not live conversation models. Flash TTS is the higher-fidelity option for character work, long-form narration, dual-speaker scenes, and line-by-line control of pacing, dialect, and acting cues.
Flash-Lite TTS is the cheaper, higher-throughput option for dubbing, read-aloud features, and voice agents. Both accept text input up to 8,192 tokens and support single-speaker and multi-speaker output, custom voice design from a written description, and voice replication from a short sample after a consent check.
Google also opened a Voices endpoint and an AI Studio audio playground where a voice can be designed, previewed, and placed in a two-speaker screenplay editor.
The voice inventory is larger than the 30 studio voices from earlier Gemini TTS releases.
Google describes more than 2,000 production-ready voices across more than 100 languages and dialects, including regional varieties such as Mexican Spanish, Quebec French, and Scots English.
Generated audio carries a SynthID watermark. Voice replication is withheld in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland, and India.
The models are available in the Gemini API, Google AI Studio, and Gemini Notebook, with Gemini Enterprise access listed as coming later. Google Vids is named as a destination for the lighter model. Hume AI’s quality ranking placed Flash TTS first and Flash-Lite TTS second on its overall index, with reported gains over Gemini 3.1 Flash TTS on long-form audio and dual-speaker control. Known limits still include occasional hallucinations, slowness, and timeouts.
Put the two side-by-side, the releases fill different gaps in the same day.
Connected Apps keep a person inside one conversation while work moves into design tools, databases, contracts, fitness catalogs, and ticket search. The 3.8 TTS models keep a developer inside the Gemini toolchain while a script becomes speech, a custom voice, or a two-person scene.
The older pattern of carrying context from window to window is not gone. It is just less automatic than it was before.























































































































































































































































































































































































