Skip to main content
The transcriber turns the caller’s audio into text. Language support is the primary selection factor — most transcribers are optimized for specific language families. Key settings: endpointing (silence detection), language (always set explicitly — auto-detection adds latency), encoding / sampling_rate (must match your telephony provider). For the exact provider and model values each takes, see Create Agent.

Next steps

Speech-to-text providers

Setup guides for every transcriber Bolna supports

Supported languages

Languages and codes you can configure

Choose an LLM

The next provider in the pipeline

Back to the decision guide

Compare all provider categories side by side