Catalog / Vercel

Gemini 3.5 Transcribe now available on AI Gateway

2 months agochangedOriginal notes

Gemini 3.5 Transcribe from Google is now available on AI Gateway for recorded and live audio:

  • google/gemini-3.5-transcribe transcribes a complete audio file in one request.

  • google/gemini-3.5-transcribe-live transcribes audio over a WebSocket and returns text as the audio arrives.

Both models automatically detect more than 85 languages, including when a speaker switches languages. You can also provide custom vocabulary to improve the transcription of names, technical terms, and uncommon spellings. The model for complete recordings can also identify speakers and return word-level timestamps.

Streaming transcription is available in AI SDK 7. Install the latest AI SDK and AI Gateway provider:

Transcribe live audio

Use streamTranscribe with a ReadableStream of raw audio chunks. Set inputAudioFormat to match the audio being sent:

Transcribe a complete recording

Use transcribe to send a complete audio file and receive the finished transcript:

Try Gemini 3.5 Transcribe Live in the model playground, browse all transcription models, or read the speech quickstart.

Read more