Skip to main content
The Languages tab (previously labelled Audio) controls how your agent listens and speaks. Pick the language you are configuring, then set its voice and transcription providers and tune audio quality. For multilingual agents, you can select different transcription and voice providers and hand-in messages per language.
Full Audio Tab view with Languages section showing English Primary, Hindi, and Dutch, Speech-to-Text with Deepgram nova-3 and Keywords, Text-to-Speech with ElevenLabs Eleven Turbo v2.5, Nila voice, and tuning sliders for Buffer Size, Speed rate, Similarity Boost, Stability, and Style Exaggeration

Audio Tab showing language selector with English as primary, Deepgram for Speech-to-Text, ElevenLabs for Text-to-Speech with voice tuning sliders


Languages

The strip at the top of the tab, labelled Settings for, selects which language the voice and transcription settings below apply to. The language marked (Default) is the one your agent opens every conversation in.
Bolna Languages Tab configuration showing English marked as the default language, with Dutch and Hindi as additional languages

Language selector showing English as the default language, with Dutch and Hindi as additional languages

Languages are added, removed, and promoted to default in the Agent Tab. On this tab the strip only switches between the languages you already configured there.

When the Agent Switches Language

For multilingual agents, a summary under the strip recaps the setup — for example “Starts in English, can move to Hindi, Dutch when …” — and ends in a dropdown that controls when a switch is allowed: requested or auto detected (the default) or the caller requested for it. See When the Agent May Switch for what each option means, and How Language Switching Works for exactly what counts as “the caller speaking another language” — why a stray English word or a read-out order number won’t trip a switch.

Supported Languages

See the Multilingual Support guide for the full list of languages Bolna can be configured with.

Voice

Controls how your agent sounds when speaking to the caller. For multilingual agents, each language can have its own voice provider, model, and voice. Select a language in the strip above to configure its voice settings independently.
Bolna Text-to-Speech settings showing Sarvam provider, Bulbul v3 model, Suhani voice selected, with sliders for Buffer Size at 230, Speed rate at 1

Text-to-Speech configuration with Sarvam as provider, Bulbul v2 as model, and Anjura voice selected, along with voice tuning sliders

Provider, Model, and Voice

1

Select a Provider

Choose from AzureTTS, Cartesia, ElevenLabs, or Sarvam.
2

Pick a Model

Select the model that fits your latency and quality needs (e.g., ElevenLabs eleven_turbo_v2_5 for low latency).
3

Choose a Voice

Click the Voice dropdown to browse all voices for the selected provider and model.

Browsing Voices

Click the Voice dropdown to see a searchable list of all available voices. Filter by gender using the All, Male, Female, and Neutral tabs. Each voice shows a play button so you can preview it before selecting.
Bolna voice selector dropdown with search bar, gender filter tabs for All Male Female and Neutral, and voices including Chef DJ, Viraj, Ben, Roger, Matt, and Angelica with play buttons

Voice selector dropdown showing a searchable list of voices with gender filter tabs and play preview buttons

Preview Welcome Message

Click the play button next to the Voice dropdown to hear the selected voice speak your agent’s welcome prompt (configured in the Agent Tab). This lets you test how the voice sounds before going live.

Voice Tuning Parameters

Fine-tune your agent’s voice using the sliders under Advanced settings, below the voice selector. Available parameters may vary by provider.
High Buffer Size improves quality but adds latency. If callers notice a delay before the agent speaks, lower this value.

Adding and Cloning Voices

Click the Add Voice + button in the Voice section to add a custom voice by ID or clone one from an audio sample.
Custom voice uploads are only available for ElevenLabs and Cartesia.

Add a Voice by ID

Use this when you already have a voice ID from your provider’s voice library.
Bolna Add Voice modal with Add by ID tab active, ElevenLabs selected as provider, Voice ID input field with placeholder pNInz6obpgDQGcFmaJgB, and a blue Add voice button

Add Voice dialog with the Add by ID tab selected, ElevenLabs as provider, and a Voice ID input field

1

Select the Add by ID tab

Make sure the Add by ID tab is selected in the dialog.
2

Choose a Provider

Select ElevenLabs or Cartesia.
3

Enter the Voice ID

Paste the voice ID from your provider. For ElevenLabs, find IDs in the ElevenLabs voice library.
4

Click Add voice

The voice will appear in the Voice dropdown for all your agents.

Clone a Voice

Create a new voice by uploading an audio recording. Useful for maintaining a consistent brand voice or using a specific person’s voice (with their permission).
Bolna Clone Voice modal with Cartesia selected as provider, Voice name placeholder Sales Assistant Voice, Description placeholder Warm male Indian accent, Sample language Hindi, and drag-and-drop upload area for audio files up to 10 MB

Clone Voice dialog with Cartesia as provider, fields for Voice name, Description, Sample language, and a file upload area

1

Select the Clone Voice tab

Switch to the Clone Voice tab in the dialog.
2

Choose a Provider

Select ElevenLabs or Cartesia.
3

Enter Voice Details

Add a Voice name (e.g., “Sales Assistant Voice”) and Description (e.g., “Warm male Indian accent”).
4

Select Sample Language

Choose the language of your audio sample.
5

Upload an Audio Sample

Drag and drop your audio file or click click to browse. Audio files only, maximum 10 MB.
6

Click Clone voice

The platform processes your sample and adds the new voice to the Voice dropdown.

Supported Languages for Voice Cloning

Both ElevenLabs and Cartesia support cloning in the same languages Bolna supports, plus an additional Indian Multilingual sample-language option.
For best results, use a clean recording with no background noise, a single speaker, and at least 30 seconds of continuous speech.

Transcription

Controls how your agent converts the caller’s spoken words into text before the LLM processes them. For multilingual agents, each language can have its own transcription provider and model. Select a language in the strip above to configure its transcription settings independently.
Different languages may perform better with different providers — see Per-Language Configuration for a worked example.
Bolna transcription settings showing Provider dropdown set to Azure, Model dropdown set to Azure, and Keywords input field with Bruce:100 as an example keyword boost entry

Transcription configuration with Azure selected as provider and model, and a Keywords field showing Bruce:100

Provider and Model

Choose a transcription provider from the Provider dropdown, then pick the specific model from the Model dropdown.

Keywords

Boost recognition accuracy for specific words the transcriber might miss, such as brand names, product names, or technical terms. Enter keywords in the format word:boost_value (e.g., Bruce:100).
Keyword boosting is only available with Deepgram. The Keywords field has no effect when using other providers.

Hand-in

Shown only when the agent has more than one language. These two fields cover how the agent moves into the language selected in the strip above.
Both are optional — leave them blank and the agent switches without a spoken hand-in line. These fields used to live under Advanced Settings in the Agent Tab, where they were called Agent Name and Handoff Message. They are still per language — the values set for Hindi do not affect English.

Next Steps

Engine Tab

Configure interruption handling, endpointing, and latency

Multilingual Support

Set up agents that speak multiple languages in a single call

How Language Switching Works

The detection engine behind mid-call switching

Clone Voices

Create a custom voice from an audio sample

Deepgram Provider

Explore Deepgram transcription models and keyword boosting