Skip to main content
OpenAI’s GPT-5 family is the current generation, with GPT-5.6 (Sol, Terra, Luna) the newest release. gpt-5.4-mini remains the default recommendation for most voice agents: it has low time-to-first-token and strong instruction following at a fraction of the cost of the full models.

Quick config

GPT-5-series models require "temperature": 1. Any other value is rejected with 400 For GPT-5 models, temperature must be 1, and the field defaults to 0.1 when omitted, so send it explicitly.
To use your own OpenAI API key, connect it at platform.bolna.ai/auth/openai.

Supported models

Recommendation: Start with gpt-5.4-mini. Step up to gpt-5.6-terra or gpt-5.6-luna for newest-generation quality at moderate cost, or gpt-5.6-sol / gpt-5.5 when you need the strongest multi-step reasoning and highest output quality (financial, medical, nuanced escalation).

Key settings

Keep max_tokens short

Voice responses should be 1–3 sentences. max_tokens: 150 is appropriate for most turns. A higher cap doesn’t hurt quality but increases tail latency on long responses. On GPT-5 models max_tokens is sent as max_completion_tokens and reasoning tokens come out of the same budget. At reasoning_effort above none/minimal, reasoning can consume most of a 150-token cap and truncate the spoken reply, so raise the cap whenever you raise the effort.

Reasoning effort

GPT-5 models reason before answering. Effort controls how much, and it is the main quality-versus-latency dial on the LLM leg of a call. Leave it unset and the model gets the lowest-latency effort it supports, which is what most voice agents want. Every model accepts a different subset, and an unsupported value is rejected when the agent is created:
minimal is valid only on gpt-5, gpt-5-mini and gpt-5-nano. On gpt-5.1 and later the equivalent is none.
For live calls, stay at none or low. Each step up adds reasoning tokens before the first spoken word, which lands directly in time-to-first-token. See Latency.

Writing prompts for voice

Prompts for voice agents differ from chat prompts:
  • Use imperative sentences: “Keep all responses under 3 sentences.”
  • Specify spoken format: “Never use bullet points or markdown — speak in complete sentences.”
  • Define handling for off-topic questions: “If asked something outside your scope, say: ‘I can only help with appointment scheduling today.’”
  • Include the welcome message in the prompt or agent config, not as part of the system prompt instructions.
See Prompting Guide for full guidance.

Function calling

All GPT-5 and GPT-4.1 models support function calling. In Bolna, functions are defined in the Tools Tab and called automatically by the LLM during conversation. gpt-5.4, gpt-5.5 and gpt-5.6 run through OpenAI’s Responses API automatically, because function calling combined with reasoning_effort is not accepted on chat completions for those models. You don’t need to configure anything for this. See Custom Function Calls for configuration.

FAQ

Use gpt-5.4-mini for most agents: lowest time-to-first-token and significantly lower cost per call. Step up to gpt-5.6-terra or gpt-5.6-luna for newest-generation quality at moderate cost, or gpt-5.6-sol / gpt-5.5 for the most demanding tasks (complex financial/medical, long multi-step tool chains) where quality is the top priority. Avoid gpt-5.5-pro for live voice; its latency is too high for real-time calls.
No. GPT-5-series models accept only temperature: 1, and anything else fails agent creation with a 400. Use the prompt to constrain behaviour instead: state the exact wording, the sentence limit, and the fallback line for off-topic questions. On the previous-generation GPT-4.1 models a lower temperature still applies.
Keep reasoning_effort at none (or minimal on gpt-5/gpt-5-mini/gpt-5-nano), lower max_tokens, use gpt-5.4-mini instead of the larger models, and write shorter system prompts (large prompts increase prefill time). Reasoning effort is usually the biggest single lever. See Latency for a full breakdown.
Yes. Connect your OpenAI account at platform.bolna.ai/auth/openai. Costs will be charged to your OpenAI account, not Bolna’s platform wallet (for the LLM component).