Skip to main content
The synthesizer turns the agent’s reply into audio. Latency and language are the primary selection factors. Always enable stream: true. Key settings: stream: true (always enable), buffer_size (100–250 chars typical), audio_format (must match your telephony provider). For the exact provider and model values each takes, see Create Agent.
gemini appears as a synthesizer provider but cannot currently be used — no Gemini TTS model resolves.

Next steps

Text-to-speech providers

Setup guides for every voice provider

Clone a voice

Create a custom voice for your agent

Choose telephony

Pick how calls reach your agent

Back to the decision guide

Compare all provider categories side by side