Gemini 3.5 Transcribe now available on AI Gateway
Gemini 3.5 Transcribe from Google is now available on AI Gateway for recorded and live audio:
google/gemini-3.5-transcribetranscribes a complete audio file in one request.google/gemini-3.5-transcribe-livetranscribes audio over a WebSocket and returns text as the audio arrives.
Both models automatically detect more than 85 languages, including when a speaker switches languages. You can also provide custom vocabulary to improve the transcription of names, technical terms, and uncommon spellings. The model for complete recordings can also identify speakers and return word-level timestamps.
Streaming transcription is available in AI SDK 7. Install the latest AI SDK and AI Gateway provider:
Transcribe live audio
Use streamTranscribe with a ReadableStream of raw audio chunks. Set inputAudioFormat to match the audio being sent:
Transcribe a complete recording
Use transcribe to send a complete audio file and receive the finished transcript:
Try Gemini 3.5 Transcribe Live in the model playground, browse all transcription models, or read the speech quickstart.