MeetStream partners: transcription engines and model providers
MeetStream partners are the speech engines and model providers you name on a request. Six transcription engines, plus the reasoning, voice and avatar providers a voice agent runs on.
*Get started in minutes*
How a provider is chosen
Name the provider when you create the bot. One runs per request. The next meeting can name another one.
*Get started in minutes*

Transcription engines MeetStream integrates
Six post-call engines. Two of them also stream live. Each card says what it does here.
Post-call transcripts on nova-3. Diarisation, smart formatting and keywords are on by default. Deepgram Streaming pushes sentence, word or raw chunks live on nova-2.
Post-call transcripts on universal-2, with speaker labels. PII redaction and chapters are options. AssemblyAI Streaming sends raw or sentence chunks in English while the call runs.
Post-call transcripts on saaras:v3. It is built for Indic languages and English code-switching. Transcribe and translate are both modes you can name on the request.
Post-call transcripts with auto-detect, or any of 162 languages named on the request. Output arrives split per speaker.
The in-house post-call engine. It spots the language on its own, translates when you ask, and needs no third-party key.
Native Google Meet and Teams captions, sent live with nothing to set up. Pair it with a post-call engine when you also want a stored transcript.
The transcription integration holds the engine table, and the transcription API covers the request itself.
Model, voice and avatar providers
A voice agent is a saved config. These are the providers you can name inside it.
Realtime models carry a whole voice-to-voice turn. In pipeline mode OpenAI is a reasoning choice. Whisper is a speech-to-text choice, and OpenAI voices read the reply out loud.
A reasoning choice in pipeline mode. It picks the reply between the speech engine and the voice.
Gemini native audio runs a realtime turn end to end. Google is also a reasoning choice in pipeline mode.
Named at every stage. It runs realtime turns, pipeline speech to text, pipeline reasoning, and pipeline voices.
A text-to-speech choice in pipeline mode, so the agent answers in the voice you pick.
Virtual avatars. Anam gives the bot a lip-synced face, set on the saved agent beside the voice.
The voice agents page sets out realtime and pipeline mode side by side.
Where each provider runs
One row per provider, and the key it runs on.
| Provider | Where it runs on MeetStream | Key |
|---|---|---|
| Deepgram | Post-call transcripts, live chunks, pipeline speech to text | Yours, or ours |
| AssemblyAI | Post-call transcripts, live chunks, pipeline speech to text | Yours, or ours |
| Sarvam | Post-call transcripts, Indic languages, translate mode | Yours, or ours |
| JigsawStack | Post-call transcripts, auto-detect across 162 languages | Yours, or ours |
| MeetStream engine | Post-call transcripts, language detection, translation | None needed |
| Meeting captions | Live captions from Google Meet and Microsoft Teams | None needed |
| OpenAI | Realtime turns, pipeline reasoning, Whisper speech to text, voices | Yours |
| Anthropic | Pipeline reasoning | Yours |
| Gemini native audio realtime, pipeline reasoning | Yours | |
| xAI | Realtime turns, pipeline speech to text, reasoning and voices | Yours |
| ElevenLabs | Pipeline text to speech | Yours |
| Anam | Virtual avatars on a saved agent | Yours |
Your keys, your accounts
Connect a provider once in the dashboard. The engine then runs on your account. Keys stay there and never travel in a request.
*Get started in minutes*

Rather run your own model? Take raw frames from the real-time audio socket and do the rest yourself.
One request, whichever provider you pick.
Speaker labels come from the stream, so they hold across every engine.
*Get started in minutes*
Got a question? We got the answer.
Common questions about which providers are available, whose key runs them, and how to swap.
*Get started in minutes*
Which transcription engines can I use?
Six post-call engines: Deepgram, AssemblyAI, Sarvam, JigsawStack, the MeetStream engine, and the meeting platform's own captions. Two live engines: Deepgram Streaming and AssemblyAI Streaming.
Which model providers can a voice agent use?
Realtime mode runs OpenAI Realtime, Google Gemini native audio or xAI. Pipeline mode splits the turn in three. Speech to text takes Deepgram, AssemblyAI, OpenAI Whisper or xAI. Reasoning takes OpenAI, Anthropic, Google or xAI. The voice takes OpenAI, ElevenLabs or xAI.
Do I need my own provider key?
Connect a Deepgram, AssemblyAI, Sarvam or JigsawStack key in the dashboard. MeetStream then runs the engine on your account and waives the $0.10 add-on. The MeetStream engine and platform captions need no third-party key.
Can I switch provider between meetings?
Yes. One engine runs per request, and the next meeting can name a different one. Re-transcribe a stored recording with a second engine to compare both on your own audio.
Do speaker labels change with the provider?
Names come from the stream each person is on, plus the speaker timeline. Switch engines and the words change while the names stay put.
Can I run my own model instead?
Yes. Subscribe to the real-time audio WebSocket. You get 48 kHz mono PCM16 frames straight from the room. Each frame carries the speaker ID and name, about 200 ms behind the talk.
Who stands behind the accuracy figures?
Accuracy and language coverage belong to each engine. Transcribe one recording twice and read both outputs to compare them on the audio you care about.
Try two engines on one recording
$5 of free credit, one engine per request, and re-transcription so you can compare them on your own audio.