Menu
See the rankings

AI voice · pricing

What an AI voice agent really costs per minute

The advertised per-minute rate is usually a floor, not a bill. Using the dated prices from our vendor pages, we split a real minute into its five parts and work four pricing shapes through with real numbers, then derive how far each platform's headline sits from the worst case.

By Voxrater · 8 min read · Published 2026-07-12

Every figure in this piece comes from the dated pricing data on our vendor pages, checked against the live vendor pricing pages between 15 June and 11 July 2026 (each capture date is in the sources below, with screenshots in our evidence store). These are estimates built from published rates, not quotes: enterprise and volume deals will land differently, and we have not yet run our own test calls.

Across the six platforms we work through below, a realistic all-in cost is roughly $0.05 to $0.31 a minute. The spread is that wide because most headline rates quote only the platform’s own fee; the parts it leaves out (the AI model, the transcription, the voice, the phone line) can take the real bill to six times the sticker.

That six-times figure is not a guess, and it is not the same for every platform. We derive it below, vendor by vendor, from the dated pricing data on our own vendor pages. Workings shown throughout.

The five parts of a minute

A voice-agent minute is really five bills stapled together. The platform fee pays the company running the call. Speech-to-text (STT) turns the caller’s words into text the software can read. The language model (the LLM, the AI brain) decides what to say back. Text-to-speech (TTS) reads that reply out in a voice. And telephony is the phone line itself. Some platforms bundle all five into one number; most quote one or two and meter the rest.

Retell itemises more cleanly than anyone, so its published rates make a good reference card:

PartThe plain-language jobRetell’s published rate (2026-07-11)
Platform / voice infrastructureRuns the call itself$0.055/min, speech-to-text included
Speech-to-textTurns the caller’s words into textincluded in the line above
Language modelThe AI brain that decides the reply$0.003 to $0.16/min, your pick from a menu
Text-to-speechReads the reply out in a voice$0.015/min ($0.040 for ElevenLabs voices)
Phone lineCarries the callabout $0.015/min on US Twilio, $0 on your own SIP line

(A SIP line just means you plug in your own phone-number supplier instead of renting the platform’s.)

The model menu moves the bill most

Look at that language-model row again, because it deserves its own table. This is the single choice that moves your bill more than the platform you pick.

Model on Retell’s menuCost per minute
GPT 5 nano$0.003
GPT 5 mini$0.012
GPT 4.1$0.045
Claude 4.6 Sonnet$0.08
Gemini 3.5 Flash$0.081
GPT 5.5$0.16 ($0.32 on the fast tier)

Top to bottom, that menu spans a 53 times spread (0.16 ÷ 0.003) inside a single platform. Two things jump out. First, the floor is genuinely cheap: 0.055 infrastructure + 0.015 voice + 0.003 nano + a free SIP line adds up to about $0.073 a minute, which matches the bottom of Retell’s own banner. Second, do not assume the famous names sort the way you expect: Gemini 3.5 Flash at $0.081 costs nearly double GPT 4.1 at $0.045. A typical build (GPT 4.1 plus a US Twilio line) works out at 0.055 + 0.015 + 0.045 + 0.015, about $0.13 a minute before add-ons.

Four pricing shapes, worked through

Every platform we track prices in one of four shapes. Here is each one with real numbers.

One flat rate. Bland bundles the model, the listening, the speaking and the phone line into a single figure: $0.14 a minute on the free entry tier, $0.12 with a $299 monthly fee, $0.11 with a $499 fee. Predictable, which is the whole appeal. The catch is arithmetic: the $299 fee saves you $0.02 a minute, so it only pays for itself past 299 ÷ 0.02, roughly 15,000 minutes a month. Below that volume, the “dearer” $0.14 tier is the cheaper bill.

A platform fee plus pass-through. Vapi charges $0.05 a minute to host the call, and that is the only number Vapi sets; the other four parts are billed at cost by whichever suppliers you wire in. Honest, but only a floor: our stored band for a realistic stacked build runs $0.05 to $0.30. ElevenLabs works the same way for its agents at a dearer base: $0.08 a minute for the premium voice, your model on top, about $0.02 for a Twilio line, landing $0.08 to $0.25. Retell, worked above, is the semi-managed version of this shape.

A bundled all-in rate. Deepgram folds speech-to-text, the model and its own voice into one $0.075 rate; add a Twilio line at about $0.014 and you get 0.075 + 0.014, about $0.089 a minute (its premium tier is about 0.163 + 0.014, call it $0.18). Speechify goes further and includes the call routing too: from $0.07 a minute, with the pricing page promising “no passthrough, no token math”. That is the counter-model to everything above, though it is a young product we have not yet put through test calls.

Subscription plus overage. Speechify’s tiers also show the fourth shape: Pro is $99 a month with 1,200 agent minutes included, then $0.07 a minute after (overage just means the metered rate once your included minutes run out). Use all 1,200 and the effective rate is 99 ÷ 1,200, about $0.0825 a minute. Use only 600 and it is 99 ÷ 600, $0.165, double the rate for the same product. The allowance you do not use is the hidden cost of this shape.

ShapeExampleThe workingsRealistic all-in
One flat rateBland$0.11 to $0.14 by tier; the cheaper rates carry $299 to $499 monthly fees$0.11 to $0.14/min
Platform fee + pass-throughVapi$0.05 platform fee + your STT, model, voice and line at cost$0.05 to $0.30/min
Platform fee + pass-throughRetell0.055 infra + 0.015 voice + model (0.003 to 0.16) + line (0 to 0.015)$0.07 to $0.31/min
Platform fee + pass-throughElevenLabs0.08 voice + your model + about 0.02 line$0.08 to $0.25/min
Bundled engine, line extraDeepgram0.075 (or 0.163 premium) + about 0.014 Twilio$0.08 to $0.18/min
Bundled incl. the lineSpeechifyone metered rate; overage $0.075 down to $0.068 by tier, Enterprise from $0.06$0.06 to $0.075/min

One more line for the budget: the add-ons stack quietly. Vapi’s HIPAA mode (the healthcare data-handling standard) is a flat $2,000 a month on top of usage, and Retell’s automatic call-quality checks cost $0.10 a minute once you pass the first hundred free minutes. Neither shows up in a per-minute comparison until you read the footnotes.

The multiples, derived

You will see claims that voice agents “really cost 2 to 4 times the advertised rate”. We can do better than a folk figure, because we store both numbers for every platform: the vendor’s headline and a sourced all-in band. Divide the top of the band by the headline and you get each platform’s honest worst-case multiple.

PlatformHeadlineOur stored all-in bandTop of band ÷ headlinePriced as of
Vapi$0.05$0.05 to $0.30/min6.0×2026-06-15
Retell$0.07$0.07 to $0.31/min4.4×2026-07-11
ElevenLabs$0.08$0.08 to $0.25/min3.1×2026-07-11
Deepgram$0.075$0.08 to $0.18/min2.4×2026-06-15
Bland$0.12$0.11 to $0.14/min1.2×2026-06-15
Speechify$0.07$0.06 to $0.075/min1.1×2026-07-11

(Deepgram’s rates were re-checked unchanged on 2026-07-11; the Retell all-in band is Retell’s own published banner, and the others are our sourced estimates of a realistic stacked build.)

Read the pattern rather than the individual numbers. The multiple is not a measure of vendor honesty; it is a measure of how much of the bill the headline includes. The pass-through platforms (6.0×, 4.4×, 3.1×) quote a floor and let your component choices set the rest. The fully bundled platforms (1.2×, 1.1×) quote nearly the whole bill, and at the bottom of their bands they come in under their own headline: Bland’s $0.11 top-volume rate sits below its $0.12 mid-tier figure, and Speechify’s $0.06 Enterprise floor undercuts its $0.07 sticker. Deepgram’s 2.4× lands in between because it bundles the engine but not the phone line, and its premium engine tier more than doubles the standard rate. So the fair summary is not “expect 2 to 4×”. It is: on a pass-through platform the worst case runs 3 to 6 times the headline and your choices decide where you land; on a fully bundled platform the headline is the bill, give or take about 15 per cent.

When each shape wins

Our position, since a table without a recommendation is just homework. If you want a number you can put in a budget and forget, buy bundled: Bland for phones at volume, Speechify if you want the cheapest bundled rate and can live with a young product and thin compliance paperwork. If you have a developer and a cost target, buy pass-through and shop the components: Retell’s $0.073 floor and Vapi’s $0.05 platform fee reward people who pick a budget model and bring their own line, with the trade-off that you now reconcile several suppliers’ bills and every saving is your own homework. If the voice itself is what sells your calls, pay ElevenLabs’ $0.08 base and accept the dearer top end; a premium voice on a bargain stack is usually the wrong way round. And if your volume is steady and predictable, a subscription with included minutes prices well, but only if you actually use the allowance you are paying for.

The honest limits

Four caveats, stated plainly. We have not yet run our own test calls, so nothing here says which platform sounds better or answers faster; this is a pricing piece, and price is one axis of several. Every figure is dated (captured between 15 June and 11 July 2026) and pricing genuinely moves: Speechify restructured its entire pricing between our June and July captures. Your prompt and model choice move the bill more than the platform choice, as the 53× menu spread shows, so treat every band here as a bracket rather than a promise. And these are estimates from published rates, not quotes; at enterprise volume most of these vendors will negotiate.

Next step: put your own monthly minutes into the cost calculator and see the six platforms above ranked for your usage, check any headline against its all-in band in the price index, or line up your two shortlist picks on the comparison pages. Each vendor page carries the full pricing notes with capture dates and screenshots.

Common questions

Why is my AI voice agent bill higher than the advertised per-minute rate?
On most platforms the advertised number covers only the platform's own fee. The AI model, the transcription, the voice and the phone line are metered on top, and on the pass-through platforms we track those parts take a realistic bill to between three and six times the headline. Bundled platforms quote much closer to the whole cost.See every platform's real range
What is the cheapest way to run an AI voice agent per minute?
At the floor, either a pass-through build with a budget model and your own phone line (Retell's published parts add up to about $0.073 a minute) or a bundled rate like Speechify's, from $0.07 with everything included. Whether that floor survives your call volume and model choice is what the calculator is for.Put your own minutes in
How much does the AI model add to the cost per minute?
More than any other component. On Retell's published menu the model alone runs from $0.003 a minute for GPT 5 nano to $0.16 for GPT 5.5, a 53 times spread inside one platform. A mid-range model such as GPT 4.1 at $0.045 roughly doubles a $0.07 voice-engine bill on its own.
Are flat-rate voice agents really flat?
Nearly. Bland's rate covers the model, the speech, the voice and the line in one number, but the cheaper rates carry monthly platform fees: $0.12 a minute costs $299 a month and $0.11 costs $499. Work the fee into your volume; it only pays for itself past roughly 15,000 minutes a month.Bland's pricing detail

Sources

Every figure above is dated and links to its primary source.

  1. Vapi pricing page (re-verified 2026-06-15, screenshot in evidence/): $0.05/min platform fee with speech-to-text, the model, the voice and telephony passed through at cost; HIPAA a $2,000/mo add-on; our stored all-in band $0.05 to $0.30/min. checked 2026-06-15
  2. Retell pricing page (captured 2026-07-11): voice infrastructure $0.055/min with speech-to-text included, text-to-speech $0.015/min (ElevenLabs voices $0.040), LLM menu from GPT 5 nano $0.003/min to GPT 5.5 $0.16/min ($0.32 fast tier), US Twilio about $0.015/min with SIP free, banner all-in $0.07 to $0.31/min, AI QA add-on $0.10/min after the first 100 minutes. checked 2026-07-11
  3. Bland pricing page (re-verified 2026-06-15, screenshot in evidence/): bundled per-minute tiers $0.14 Start ($0/mo), $0.12 Build ($299/mo), $0.11 Scale ($499/mo), each covering the model, speech-to-text, the voice and the phone line in one rate. checked 2026-06-15
  4. ElevenLabs Agents pricing page: $0.08/min base call rate on every plan tier, with the language model billed separately as pass-through. checked 2026-06-02
  5. ElevenLabs pricing page re-captured 2026-07-11: the $0.08/min agents rate unchanged; our stored all-in band is $0.08 to $0.25/min once the model and a Twilio line (about $0.02/min) are added. checked 2026-07-11
  6. SpeechifyAI pricing page (captured 2026-07-11): Voice Agents from $0.07/min all-in ('no passthrough, no token math'), covering the model, speech-to-text, the voice and telephony orchestration; tier allowances of 60, 120, 1,200 and 6,000 minutes with overage $0.075, $0.07 and $0.068 by tier; Enterprise from $0.06/min. checked 2026-07-11
  7. Deepgram pricing page (re-verified 2026-07-11, rates unchanged): Standard Voice Agent $0.075/min bundling speech-to-text, the language model and Aura-2 text-to-speech; Advanced about $0.163/min; telephony not included, brought via Twilio at about $0.014/min (a third-party cost). checked 2026-07-11

Get the next piece

New analysis and dated test results land in the newsletter first. No spam.

Newsletter launching soon.