What a voice agent actually costs per minute
Four line items stack into every minute of an AI phone call: hosting, model tokens, speech conversion, and the phone line itself.

The short answer
- A voice agent's per-minute cost has four layers: hosting and orchestration, LLM tokens (priced by the model), text-to-speech (per character) and speech-to-text (per minute). Telephony sits on top, priced per minute with regional differences.
- A quoted 'all-in' price is only comparable to another vendor's if both bundle the same layers. A price that leaves out telephony, or assumes one model tier, looks cheaper without being cheaper.
- Model choice sets both the token cost and the latency of a call. A call tolerates a slow or wrong response less than a text chat, because the caller waits in silence.
- SalesCrew's outbound voice calls run through Vapi. A per-minute calculator that publishes each layer's dated vendor rate compares voice agent products better than one blended number.
Why 'per minute' hides four different prices
A vendor quoting one number per minute is usually bundling four costs from four separate providers. Unbundling them is the only way to compare two vendors, or to estimate what a call will cost before running it.
The first layer is hosting and orchestration. It runs the call session, manages the conversation state, and coordinates the other layers in real time. The second is the language model that generates the agent's responses. It is priced by input and output tokens at the model provider's own rate, which varies by model and by how much the agent talks versus listens. The third is text-to-speech, which turns the model's text into audio, usually priced per character. The fourth is speech-to-text, which turns the caller's audio into text the model can read, usually priced per minute of audio. On top of all four sits telephony: the phone line itself. It is priced per minute, and the rate depends on the carrier, the call direction, and the region.
Where the real cost differences come from
Model choice has the widest cost range. Providers price by token, and the rate can differ by ten times or more between a fast, low-cost model and a slower, higher-quality one. A voice agent also has a latency budget that a chat interface does not. A caller sitting in silence for two seconds notices. A chat user waiting for a reply usually does not. So the model choice trades cost, speed and answer quality at the same time, not cost alone.
Text-to-speech and speech-to-text costs scale with how much is said and how much is heard. A talkative caller or a verbose agent costs more in those two layers than a short exchange, whatever model is used. Some vendors also sell voice quality in tiers at different prices. That is another place where two "per minute" quotes can diverge without either being wrong.
Telephony is the layer least connected to the AI stack, and the one most often under-quoted. A phone carrier prices it on its own schedule. Rates differ by country, by whether a number is toll-free, and by call direction. A vendor quoting an "all-in" rate has built in one telephony assumption. That assumption may not match your calling pattern.
How to actually estimate or compare a per-minute cost
The formula is plain. Cost per minute equals the hosting fee, plus LLM input tokens per minute times the input price, plus output tokens per minute times the output price, plus text-to-speech characters per minute times the per-character price, plus speech-to-text minutes times the per-minute price, plus the telephony rate per minute. Every input has a dated, published rate from its provider: the model's pricing page, the speech vendors' pricing pages, and the carrier's rate card. So the formula can be filled with real, sourced numbers rather than a vendor's blended claim.
The practical use is honest comparison. Ask what each layer uses. Get each layer's own rate. Then compute the total for your expected call pattern: average call length, how much the agent talks versus the caller, and the telephony region. Do not trust one quoted number to mean the same thing across two vendors. SalesCrew's outbound calling runs through Vapi today. A calculator that publishes each layer's rate with a source and a date is a better tool than any single number on this page, because provider rates change and a static number goes stale.
Questions
- Why do two voice agent vendors quote such different per-minute prices?
- Usually because they bundle different sets of the four cost layers, or use different providers for each. A quote that only covers the platform's markup on top of pass-through provider costs looks cheaper than one that bundles everything. The gap closes once the pass-through costs are added.
- Is a cheaper LLM always the right way to cut voice agent cost?
- Not automatically. A faster, cheaper model cuts latency and cost per minute. But a voice call tolerates a wrong or confused answer less than a chat does, because the caller cannot scroll back to reread. The trade is cost against the model's reliability on the task, not cost alone.
- Does telephony cost scale differently from the AI layers?
- Yes. The AI layers (the model, text-to-speech and speech-to-text) are priced per minute or per unit of text or audio, whatever the call volume. Telephony carriers use volume pricing and regional rates. The same call can cost different telephony fees depending on where it starts or ends.