Every rate in this piece comes from the dated pricing captures on our vendor pages, checked against the live vendor pages between 30 May and 12 July 2026 (capture dates in the sources below, with screenshots in our evidence store for the July sweep). The characters-per-minute key is our documented editorial basis, not a vendor figure, and every derived rate is labelled as derived where it appears.
Every voice AI price converts to one number: cost per finished minute of speech. The master key is that a finished minute comes out of roughly 900–1,000 characters of script, so a 10-minute job is about 9,000 characters. Convert each vendor’s unit to that and the wrappers stop mattering.
That paragraph is the whole decoder; the rest of this page applies it, one unit per section, each section ending with what the identical 10-minute job costs in that wrapper. Each unit is a currency the vendor invented; you want the exchange rate.
The master key: 900 to 1,000 characters per minute
Speech generators bill by the text you feed them, so the key conversion is characters to minutes. Our documented basis, stated on the Rime page and in the narration roundup: 1,000 characters of script is roughly 150–180 words, and a finished minute of speech comes out of roughly 900–1,000 characters. So a 10-minute script is about 9,000 characters. On a live call the agent only speaks about half the time, so a conversation minute burns roughly 450–500 characters of generation.
Two honesty notes. The key is our editorial basis, not a vendor’s, and a vendor’s own conversion can disagree (Fish Audio’s does; we work the difference through below). And it prices generation only; the tier carrying your commercial licence is often the bigger line.
Credits: ElevenLabs, Cartesia, Fish Audio
A credit is a prepaid balance you spend as you generate, and it means whatever the vendor says it means.
ElevenLabs keeps it clean: on the standard models one character costs one credit, and Flash bills at half the per-character rate by its own docs. The API prices the same units in money: $0.10 per 1,000 characters on Multilingual, $0.05 on Flash. On subscription the effective rate depends on the tier, and only holds if you use the whole allowance:
| ElevenLabs tier | Price per month | Credits included | Effective rate per 1,000 |
|---|---|---|---|
| Starter | $6 | 30,000 | $0.20 |
| Creator | $22 | 121,000 | about $0.18 |
| Pro | $99 | 600,000 | about $0.165 |
Cartesia meters by time instead: about 15 credits per second of generated audio, no flat per-character rate any more. Pro is $5 a month for 100,000 credits, Startup $49 for 1.25 million, Scale $299 for 8 million.
Fish Audio’s subscriptions are credits too (Plus $15 for 250,000, Pro $100 for 2 million), and its plan page cannot agree with itself on what one buys: the FAQ says a minute costs 600–625 credits, while the plan cards imply roughly 1,250 (4,000 on the top tier). Same page, same capture, 11 July 2026.
The 10-minute job, in credits. ElevenLabs: 9,000 characters = 9,000 credits, so at the API’s $0.10 per 1,000 that is $0.90 (Flash: $0.45). Cartesia: 10 minutes = 600 seconds × 15 = 9,000 credits, costing $0.45 on Pro ($5 ÷ 100,000 × 9,000), about $0.35 on Startup and $0.34 on Scale (our stored derived $0.035 per 1,000 matches those two tiers; Pro runs nearer $0.05). Fish on its FAQ’s conversion: 6,000–6,250 credits, about $0.36–0.38 on Plus; on its cards’ implied conversion: 12,500 credits, about $0.75. One page, one job, double the price.
Characters, billed as characters
Speechify is the simple case: Starter at $10 a month includes a million characters, then $10 per million after, which is $0.01 per 1,000; the overage rate falls to $8 per million on Pro and $6 on Scale. Hume meters its Octave engine at $0.15 per 1,000 characters on the entry tiers, falling to $0.05 on the $500-a-month Business plan. Rime publishes a single “starting at” rate of $0.05 per 1,000 characters, with volume discounts it no longer spells out. And Telnyx’s cheapest listed voice is $0.000009 per character, which looks alarming and is $0.009 per 1,000.
The 10-minute job, in characters. Speechify: 9 × $0.01 = $0.09 (about $0.054 at Scale overage). Hume: 9 × $0.15 = $1.35 at entry ($0.45 on Business). Rime: 9 × $0.05 = $0.45, a starting-at figure. Telnyx floor: 9,000 × $0.000009 = about $0.08; premium voices cost more per character.
Bytes: Fish Audio’s API
Fish’s API does not bill characters at all: it bills $15.00 per million UTF-8 bytes on every current model. A byte is the smallest unit a computer stores, and in the UTF-8 text encoding an English letter or space takes exactly one byte, so for English scripts bytes and characters are the same number. Other scripts are not so lucky: under the encoding standard (RFC 3629), characters outside the basic Latin set take 2–4 bytes each, so a Chinese or Japanese script at three bytes per character costs roughly three times the English figure. That is the encoding standard, not a Fish surcharge, but it lands on the bill.
The 10-minute job, in bytes, twice. By our key: 9,000 English characters is about 9,000 bytes, so 9,000 ÷ 1,000,000 × $15 = about $0.14. By Fish’s own conversion (“1M UTF-8 bytes is approximately 180,000 English words, or about 12 hours of speech”): $15 ÷ 720 minutes ≈ $0.021 per minute, so 10 minutes = about $0.21. The two disagree because Fish’s conversion implies about 1,389 bytes per spoken minute (1,000,000 ÷ 720) against our 900–1,000; neither is wrong, they just assume different speech rates. Budget on the $0.21 and be pleased if your script comes in nearer $0.14.
Bundled hours: Murf
Murf sells finished audio by the year. Creator is $19 a month billed annually, which is $228 a year for 24 hours of generation; Business is $66 a month, $792 a year, for 96 hours. Murf publishes no per-character or per-minute rate, so the division is ours: $228 ÷ 1,440 minutes ≈ $0.16 a minute on Creator, $792 ÷ 5,760 ≈ $0.14 on Business. The shape matters as much as the rate: bundled hours are pre-bought, so an allowance you do not use is money already spent.
The 10-minute job, in bundled hours. Creator: 10 × $0.158 ≈ $1.58. Business: 10 × $0.1375 ≈ $1.38. Our narration roundup prints $1.44 for the same job because it derives per 1,000 characters ($0.16 × 9) rather than per bundled minute. Both are honest divisions of the same $228; the gap is rounding inside the characters-per-minute key, so we show both rather than average them.
Per minute, with and without passthrough: live calls
Live call platforms quote per minute, and the question is what the minute includes. Vapi charges a $0.05-a-minute platform fee and passes the rest through: speech-to-text, the model, the voice and the phone line billed at cost from the providers you wire in (passthrough means no markup on those parts). Our recorded all-in band for a realistic stacked build is $0.05–0.30 a minute. Retell’s own banner says $0.07–0.31 all-in, and its published floor adds up: $0.055 infrastructure + $0.015 voice + $0.003 cheapest model + free SIP ≈ $0.073. Speechify’s agents go the other way: from $0.07 a minute all-in, no passthrough, no token maths, says the pricing page.
One warning: a call minute is not a narration minute. The price covers listening, reasoning, speaking and the phone line while both sides talk, so never compare it against a per-character rate. We dissect the call minute in what a voice agent really costs per minute.
The 10-minute job, as a call. Vapi: 10 × $0.05–0.30 = $0.50–3.00, your stack decides. Retell: 10 × $0.07–0.31 = $0.70–3.10 (a typical GPT 4.1 and Twilio build runs about $0.13 a minute, so $1.30). Speechify agents: 10 × $0.07 = $0.70 flat.
Tokens: OpenAI’s Realtime API
A token is the small chunk of data an AI model reads and writes, and OpenAI Realtime bills voice in them. The conversion is vendor-documented rather than folklore: OpenAI’s own cost guide states that input audio costs 1 token per 100ms and output audio 1 token per 50ms, so 600 input tokens and 1,200 output tokens per spoken minute. The audio rates are $32 per million input tokens ($0.40 cached) and $64 per million output tokens.
The 10-minute job, in tokens (a call, half caller and half agent). Five minutes of caller audio in: 5 × 600 = 3,000 tokens, × $32 per million ≈ $0.10. Five minutes of agent audio out: 5 × 1,200 = 6,000 tokens, × $64 per million ≈ $0.38. Raw audio total: about $0.48. And honestly, the floor is not the bill: your instructions plus the conversation so far are re-sent as text context on every turn, and that usually dominates. One independent analysis we track puts the realistic range at $0.18–0.46 a minute uncached and $0.05–0.10 with prompt caching, so $1.80–4.60 or $0.50–1.00 for the call; the analyst’s model, not an OpenAI rate.
Hours of audio: the transcription side
Speech-to-text vendors price the listening, usually per hour of audio, and the model labels matter. AssemblyAI’s real-time streaming starts at $0.15 an hour on Universal-Streaming; its asynchronous transcription (after-the-fact, not live) runs $0.15 an hour on Universal-2 and $0.21 on Universal-3.5 Pro; and the premium real-time model, Universal-3.5 Pro Realtime, is $0.45 an hour. Quote “$0.45 streaming” unqualified and you overstate the entry price three times. Deepgram meters the same job per minute instead, and Speechmatics shows discounted rates by default (more in the traps below).
| Vendor | Unit | Base real-time rate | The 10-minute job |
|---|---|---|---|
| AssemblyAI | per hour | $0.15/hr ($0.0025/min) | $0.025 (premium model: $0.075) |
| Speechmatics | per hour | $0.24/hr as displayed, discount ON | $0.04 (undiscounted runs higher) |
| Deepgram | per minute | $0.0048/min Nova-3 streaming | $0.048 |
The 10-minute job, transcribed. 10 × $0.0025 = $0.025 on AssemblyAI’s base streaming ($0.075 premium); 10 × $0.0048 = $0.048 on Deepgram Nova-3. Listen-only costs: a whole agent also reasons and speaks, so a $0.004 transcription minute and a $0.07 agent minute are not rivals.
The master conversion table
The same 10-minute job in every unit, each figure worked in its section. The call and transcription rows price a different job from the narration rows: a unit decoder, not a ranking.
| Unit | Who uses it | What the 10-minute job costs |
|---|---|---|
| Credits | ElevenLabs (1 character = 1 credit), Cartesia (about 15/second), Fish Audio plans | $0.90 ElevenLabs standard ($0.45 Flash); $0.34–0.45 Cartesia by tier; $0.36–0.75 Fish, conversion-dependent |
| Characters | Speechify, Hume, Rime, Telnyx | $0.09 Speechify; about $0.08 Telnyx floor; $0.45 Rime starting-at; $1.35 Hume entry |
| UTF-8 bytes | Fish Audio’s API | about $0.14 by our key; about $0.21 by Fish’s own conversion |
| Bundled hours | Murf studio plans | about $1.58 Creator, $1.38 Business (our division) |
| Per minute (call) | Vapi, Retell, Speechify agents | $0.50–3.00 Vapi; $0.70–3.10 Retell; $0.70 Speechify |
| Tokens | OpenAI Realtime | about $0.48 raw audio; $0.50–4.60 realistic with text context |
| Hours of audio (STT) | AssemblyAI, Speechmatics; Deepgram per minute | $0.025–0.075 AssemblyAI; $0.04 Speechmatics displayed; $0.048 Deepgram (listen-only) |
The traps our captures keep catching
Six patterns, all caught on live pricing pages, all with dated screenshots.
The page that disagrees with itself. Fish’s FAQ says 600–625 credits a minute; its own plan cards imply roughly 1,250. When a page self-contradicts, price on the dearer reading.
The default-ON discount toggle. Speechmatics’ rate table renders with a “Model Training” toggle enabled, a 33 per cent discount in exchange for your audio training its models. The displayed price assumes a data deal you have not agreed to. Toggle it off before quoting.
The promo price where the regular price belongs. ElevenLabs’ Creator card showed $11 at our 11 July capture, a first-month offer; the regular price is $22. We briefly recorded the promo as the base ourselves, which is rather the point: it catches people who compare prices for a living.
Annual pricing dressed as monthly. Murf’s “$19 a month” Creator plan bills annually, a $228 commitment up front. Cartesia and Fish have shown annual-equivalent monthly figures too. Check for “billed annually” near the big number.
The geo-priced page. Fish’s plan page rendered in pounds from a UK connection until we switched its currency picker to USD. Compare it against a dollar page unaware and you are off by the whole exchange rate.
The starting-at rate. Rime’s “STARTING AT $0.05/1K characters” comes with unpublished volume discounts and placeholder FAQ text still live at our capture. A fine floor, a poor budget. Get the rate card in writing.
How to compare any two platforms in three steps
Step one: find the real unit price. Not the plan price, the per-unit price: dollars per 1,000 characters, per million bytes, per minute or per hour. If the vendor only sells bundles, divide the price by the units inside and call the result your own division, as we do with Murf.
Step two: convert to cost per finished minute. Use the key: 900–1,000 characters per finished minute, about half that per conversation minute for a live agent. If the vendor publishes its own conversion, run both, as we did for Fish; when they disagree, budget on the dearer one.
Step three: price your real month, not the worked example. Multiply by your actual volume, then add the subscription floor and the allowance you will not use. This is where bundled hours and use-it-or-lose-it credits quietly change the answer.
Our calculator runs all three steps against every platform we track, using the stored, dated rates behind this page. Put your own volume in and let it argue with your shortlist.
Common questions
How do ElevenLabs credits work?
How many characters is one minute of AI speech?
Why do prices on the same vendor page disagree with each other?
Is a per-minute call rate comparable with a per-character narration rate?
Sources
Every figure above is dated and links to its primary source.
- ElevenLabs pricing page (captured 2026-07-11, screenshot in evidence/): 1 character = 1 credit on the standard models; tier allowances Free 10,000 credits, Starter $6/30,000, Creator $22/121,000, Pro $99/600,000, Scale $299/1.8M, Business $990/6M; the Creator card displayed an $11 first-month promo against the $22 regular price. checked 2026-07-11
- ElevenLabs API pricing page (fresh capture 2026-07-12): Multilingual v2/v3 at $0.10 per 1,000 characters, Flash at $0.05 per 1,000. checked 2026-07-12
- ElevenLabs models docs (fresh capture 2026-07-12): Flash billed at a 50% lower price per character for API generations. checked 2026-07-12
- Cartesia pricing page (captured 2026-07-11, screenshot in evidence/): TTS metered in credits at about 15 credits per second of generated audio, with no flat per-character rate published; Pro $5/mo with 100,000 credits, Startup $49 with 1.25M, Scale $299 with 8M. checked 2026-07-11
- Fish Audio plan page (captured 2026-07-11 after switching its currency picker from geo-rendered GBP to USD): subscription credits Free 8,000/mo, Plus $15/250,000, Pro $100/2M, Max $999/25M; the FAQ states 600 to 625 credits per minute while the plan cards imply about 1,250 (Max: 4,000), on the same page. checked 2026-07-11
- Fish Audio API pricing docs (re-confirmed 2026-07-12): $15.00 per 1M UTF-8 bytes on every current TTS model, with the vendor's own conversion '1M UTF-8 bytes is approximately 180,000 English words, or about 12 hours of speech'; Fish STT at $0.36 per audio hour. checked 2026-07-11
- RFC 3629, the UTF-8 encoding standard (captured 2026-07-12): ASCII characters encode as 1 byte; characters outside the basic Latin range take 2 to 4 bytes each. checked 2026-07-12
- SpeechifyAI pricing page (captured 2026-07-11, screenshot in evidence/): Starter $10/mo including 1M characters then $10 per 1M, Pro $99 including 3M then $8 per 1M, Scale $499 including 10M then $6 per 1M; Voice Agents from $0.07/min all-in with tier overage down to $0.068 and Enterprise from $0.06. checked 2026-07-11
- Hume pricing page (re-captured 2026-06-15): Octave TTS at $0.15 per 1,000 characters on the entry tiers, down to $0.05 per 1,000 on the $500/mo Business plan. checked 2026-06-15
- Rime pricing page (captured 2026-07-11, screenshot in evidence/): a single 'STARTING AT $0.05/1K characters' rate with volume discounts unpublished; the earlier per-model rates (Mist at $0.03) removed in the mid-2026 redesign, with placeholder FAQ text still on the page. checked 2026-07-11
- Telnyx voice AI page (captured 2026-07-11): cheapest listed per-character voice (Amazon Polly standard) at $0.000009 per character, $9 per million. checked 2026-07-11
- Murf pricing page (captured 2026-05-30, tiers re-checked 2026-06-15): Creator $19/mo billed annually ($228/yr) for 24 hours of generation a year; Business $66/mo ($792/yr) for 96 hours a year; no flat per-minute or per-character rate published. checked 2026-05-30
- Vapi pricing page (re-verified 2026-06-15, screenshot in evidence/): $0.05/min platform fee with speech-to-text, the model, the voice and telephony passed through at cost; our stored all-in band is $0.05 to $0.30/min. checked 2026-06-15
- Retell pricing page (captured 2026-07-11, screenshot in evidence/): banner all-in $0.07 to $0.31/min; published components $0.055/min voice infrastructure with speech-to-text included, text-to-speech $0.015/min, GPT 5 nano $0.003/min, SIP free. checked 2026-07-11
- OpenAI API pricing (captured 2026-07-11): gpt-realtime audio tokens at $32.00 per 1M input, $0.40 per 1M cached input, $64.00 per 1M output. checked 2026-07-11
- OpenAI realtime cost guide (fresh capture 2026-07-12): input audio costs 1 token per 100ms and output audio 1 token per 50ms, which is 600 input tokens and 1,200 output tokens per spoken minute. checked 2026-07-12
- AssemblyAI pricing page (re-verified 2026-07-11): Universal-Streaming real-time at $0.15/hr; asynchronous Universal-2 at $0.15/hr and Universal-3.5 Pro at $0.21/hr; the premium Universal-3.5 Pro Realtime model at $0.45/hr. checked 2026-07-11
- Deepgram pricing page (re-verified 2026-07-11): speech-to-text metered per minute, Nova-3 streaming at $0.0048/min. checked 2026-07-11
- Speechmatics pricing page (captured 2026-07-11, JS-rendered table read from the dated screenshot in evidence/): real-time Standard at $0.24/hr, displayed with the 'Model Training' 33% discount toggle ON by default, so the shown rates are the discounted ones. checked 2026-07-11
- Independent (secondary) cost analysis of OpenAI Realtime across 11 call profiles: a realistic $0.18 to $0.46/min uncached and $0.05 to $0.10/min with prompt caching. An analyst's model, not an OpenAI rate. checked 2026-05-31
Get the next piece
New analysis and dated test results land in the newsletter first. No spam.
Newsletter launching soon.