Pricing and the model line-up were read from ElevenLabs' own pricing, models docs and v3 announcement on 18 June 2026, and both move, so treat the figures as dated rather than permanent. The ~75ms latency is ElevenLabs' own number, not a Voxrater measurement, and the quality calls here are our editorial read, not a blind listening test (we have not run one yet).
ElevenLabs sells four voice models, and the job decides which one. For narration and video, use v3 for the most expression or Multilingual v2 for proven range. For a live phone agent, use Flash v2.5: it answers in about 75ms and costs half the credits. Turbo is now retired.
Paid link, we may earn a commission. How this works.
So that is the short answer. Here is the reasoning, the real cost per job, and where each model wins.
A quick plain-language note first. A voice model is the engine that turns your text into speech (text-to-speech, or TTS). ElevenLabs offers several, and newer or faster ones often cost less per character. You pay in credits, a prepaid balance you spend as you generate, and on the standard models one character of text equals one credit.
Four models, and one has just been retired
The line-up is not a simple ladder from worst to best. The models split into two camps. v3 and Multilingual v2 are the quality models, made for speech you record once and play many times. Flash v2.5 is the speed model, made to answer a real person in real time. They are different tools for different jobs, so “which is best” only makes sense once you say what you are doing.
The big change this year: Turbo v2.5 is gone. ElevenLabs’ own model list no longer shows Turbo and notes that the Flash models replaced it at the same low latency. So if you were weighing “v3 vs v2.5 vs v2”, the live three-way is really v3 vs Flash v2.5 vs Multilingual v2.
| Model | Best for | Languages | Credits per character | Latency |
|---|---|---|---|---|
| Eleven v3 | Most expressive narration and dubbing | 70+ | standard rate | not built for real time |
| Multilingual v2 | Audiobooks and video narration | 29 | 1 | not built for real time |
| Flash v2.5 | Live phone agents and real time | 32 | 0.5 | ~75ms (ElevenLabs’ figure) |
| Flash v2 | English-only real time | English only | ~0.5 | ~75ms (ElevenLabs’ figure) |
| Turbo v2.5 | Retired, use Flash v2.5 | n/a | n/a | n/a |
ElevenLabs publishes a half-price-per-character rate for Flash but does not separately break out v3’s credit cost, so we treat v3 at the standard one-credit rate, the same as Multilingual v2. Read the latency figure as the vendor’s claim, not our measurement: we have not yet put these models through our own test calls.
What each one actually costs
Cost comes down to two numbers: how many credits a character burns, and how many credits your plan includes. Flash is the cheap one because it bills at about 0.5 credits per character against 1 credit on v3 and Multilingual v2. Same words, half the credits. In rough money, a thousand characters runs about $0.09 to $0.20 of your monthly credit depending on tier, and roughly half that on Flash.
Worked through two real jobs (assuming about 150 spoken words a minute, so roughly 9,000 characters in a ten-minute script):
| Job | Rough size | v3 or Multilingual v2 | Flash v2.5 |
|---|---|---|---|
| A 10-minute YouTube voiceover | ~9,000 characters | ~9,000 credits | ~4,500 credits |
| One audiobook chapter | ~50,000 characters | ~50,000 credits | ~25,000 credits |
| Creator plan ($22/mo) covers | 121,000 credits/mo | ~13 of those scripts | ~26 of those scripts |
A live phone agent is metered differently. There you pay ElevenLabs’ agent rate of $0.08 a minute (the same on every plan), with the language model billed separately on top, so the model choice is about latency, not per-character cost. The free plan’s 10,000 credits cover about ten minutes of narration or fifteen minutes of agent time, with no commercial licence, so anything you sell or publish needs Starter at $6 a month or above.
Want this in your own numbers? Put your monthly volume into our cost calculator, or check the headline rate against every platform in the price index.
Which model for which job
This is where we take a position. The honest read, with quality treated as our editorial judgement rather than a scored listening test:
| Your job | Use this | Why |
|---|---|---|
| Live phone agent (sales, support) | Flash v2.5 | Lowest latency, half the credits, the model ElevenLabs ships for agents |
| YouTube, ads, dubbing | Eleven v3 | The most expressive model, and listeners hear it repeatedly |
| Audiobooks and long narration | Multilingual v2 | Proven, lifelike, and built for long-form |
| A language v3 does not cover well | Multilingual v2 or Flash v2.5 | Check the per-model language list before you commit |
| Tightest budget | Flash v2.5 | Half the credits per character, so your plan stretches twice as far |
For narration and video: v3, or Multilingual v2
If the audio is recorded once and played many times, quality is worth the slower render. v3 is our pick for expression. It is ElevenLabs’ newest model, now out of alpha and generally available, and it reads emotional cues and tone shifts more naturally than anything before it, with “audio tags” you can drop into the text to direct a laugh or a sigh. It also covers the widest language set, 70+ against Multilingual v2’s 29.
Reach for Multilingual v2 instead when you want a model with a long track record for audiobooks and steady long-form narration, or when its 29 languages already cover you. Both bill at the standard credit rate, so the choice is about sound, not cost.
For a live phone agent: Flash v2.5, no argument
A phone call is unforgiving: every extra hundred milliseconds of silence reads as a stall. Flash v2.5 is the only sensible choice here. ElevenLabs quotes about 75ms to first audio, recommends it for real-time and its Agents platform, and prices it at half the credits. Just as important, v3 cannot run in real time, by ElevenLabs’ own guidance, so do not be tempted to put the prettier model on a live line. Flash v2 is the same idea if you are English-only.
The honest limits
Three things to keep in mind. The quality ranking above is our editorial read, not a blind listening test, so treat “most expressive” as a considered opinion until we publish measured results. The ~75ms latency is ElevenLabs’ figure, measured excluding your app and network, so your real-world number will be higher. And pricing plus the model list both move quickly, which is why every figure here is dated to 18 June 2026 and linked to its source.
For where ElevenLabs sits against the alternatives, see our ElevenLabs review, or put it head to head with Vapi or Retell. If voice quality is the whole point of your project, it is still the one to beat.
Paid link, we may earn a commission. How this works.
Common questions
Which ElevenLabs model should I use for a live phone agent?
Is Eleven v3 worth paying for over Flash?
Which ElevenLabs model is cheapest?
What happened to Turbo v2.5?
Can I use Eleven v3 on the free plan?
Sources
Every figure above is dated and links to its primary source.
- ElevenLabs models docs: the live text-to-speech line-up (Eleven v3, Multilingual v2, Flash v2.5, Flash v2), per-model languages and character limits, Flash at ~75ms, and the recommendation to use Flash v2.5 for real-time agents. Turbo no longer listed. checked 2026-06-18
- ElevenLabs pricing: subscription tiers and monthly credits (Free 10k to Business 6M), 1 character = 1 credit on the standard models, Flash/Turbo at 0.5 to 1 credit per character, commercial licence from Starter ($6/mo). checked 2026-06-18
- ElevenLabs Agents pricing: $0.08 per minute base call rate, the same on every plan, with the language model billed separately, plus included agent minutes per tier. checked 2026-06-18
- ElevenLabs: Eleven v3 is now generally available (out of alpha), with a 68% drop in complex-text errors over the earlier release. checked 2026-06-18
Get the next piece
New analysis and dated test results land in the newsletter first. No spam.
Newsletter launching soon.