Google has made Gemini 3.5 Transcribe generally available as a dedicated speech-to-text surface in the Gemini ecosystem. The non-streaming model automatically detects more than 85 languages, supports speaker diarization and word-level timestamps, accepts custom vocabulary hints and can clean up disfluencies and formatting through Google’s Smart transcription mode. A separate Transcribe Live endpoint handles low-latency streaming over the Gemini Live API.
The release is more specialized than the general Gemini assistant. Google documents limits that matter for comparison: file transcription supports up to one hour of audio, reduced to 30 minutes when features such as diarization or word-level timestamps are enabled; live sessions are limited to 10 minutes, and the live endpoint does not provide diarization or word-level timestamps. Google has also exposed the models through the Gemini API and AI Studio, while related dictation features are appearing in Gemini product surfaces.
For AiToolMap users, the important point is product-surface matching. This is a dedicated transcription capability within the Gemini ecosystem, not evidence that the ordinary Gemini chat interface should automatically be ranked against specialist transcription products on every workflow. The release nevertheless materially broadens Gemini’s audio tooling and warrants a revalidation of the current Gemini review.