Fish Audio
Expressive text-to-speech and 15-second voice cloning for produced audio, with an API at $15 per million characters.
Paid link, we may earn a commission. How this works.
Scored on the same voice-agent rubric as the full platforms, so a building block like this scores low on the axes it does not address. Read its value score against its job.
See how it stacks up · Full rankings →Expressive AI speech for voiceovers, audiobooks and character audio. Pick from a 2-million-voice community library or clone a voice from a short clip. The developer rate works out near $0.015 per 1,000 English characters. Not a phone-agent platform.
About $0.02 to 0.07 for a minute of finished voiceover. Narration is billed by how much text you turn into speech, not by call time.
That's roughly $1.26–4.50 an hour of audio. Plans: from $0/mo (Free) up to $999/mo (Max).
Pricing
Show the cost breakdown
| What the platform charges to run the agent, before the phone line and the AI usage are added on. | — |
|---|---|
| The step that turns what the caller says out loud into text the AI can read. | — |
| The AI 'brain' that reads what the caller said and works out what to say back. | — |
| The step that turns the AI's written reply back into a spoken voice. | $0.02 /min |
| The phone line itself: the service that connects the call to a real phone number. Usually billed on top of the platform. | — |
| The total you actually pay for one minute of conversation once every piece is added up: the platform, the AI, the voice and the phone line. | $0.02–0.07 /min |
Fish Audio is a produced-audio TTS vendor: no phone product, no telephony components, so the per-minute figure here is a DERIVED effective cost of generated audio, not a call rate. Workings: the API charges $15 per 1M UTF-8 bytes on every current TTS model (s1, s2-pro, s2.1-pro), and the docs' own conversion says 1M bytes is about 12 hours of speech, so $15 over 720 minutes is about $0.021 per generated minute (the low end and the headline). The top of the band is the entry paid tier's effective rate: Plus at $15 list for up to 200 minutes is $0.075/min; Pro sits between at about $0.062. The plan page contradicts itself on credits per minute: the FAQ says 600 to 625, the Plus and Pro cards imply about 1,250 (250,000/200 and 2,000,000/1,620), and the Max card implies 4,000 (25,000,000/6,250). We record the cards as printed and exclude Max's implied $0.16/min from the band because its printed minutes contradict the same page twice over; no tier ladder prices its biggest plan above its entry tier. At capture the page ran a countdown promotion (annual billing at 3 months off plus a holiday 50% off: $66/$450/$8,988), so treat annual figures as promotional. The page geo-prices (it rendered GBP from a UK connection); we captured after switching the page's own currency picker to USD (ADR-0010). Pay-as-you-go API access requires a paid subscription; a free s2.1-pro-free API model exists for development with fair-use limits. Speech-to-text (transcribe-1) is $0.36 per audio hour; Voice Design is $0.01 per successful request; API concurrency scales with lifetime spend (5 parallel requests below $100, 15 from $100, 50 from $1,000).
Every plan in one place: the monthly fee, what each one includes, and the features it unlocks. Anything beyond a plan's allowance, or on a pay-as-you-go tier, is billed at the per-minute rate above. A blank in the features means the vendor's plan page does not state it for that plan, not that it is unavailable.
| Free | Plus | Pro | Max | Enterprise | |
|---|---|---|---|---|---|
| Price | Free | $15/mo | $100/mo | $999/mo | Custom |
| Included | 8,000 credits ≈ 9 min of audio | 250,000 credits ≈ 4.6 hrs of audio | 2,000,000 credits ≈ 37 hrs of audio | 25,000,000 credits ≈ 463 hrs of audio | — |
| Plan notes | Card sizes it at up to 7 minutes of generation; 500 characters per generation; 3 public voice slots; no card needed. Voice cloning and commercial use are crossed out on the card: free output is personal use only (the FAQ and the terms both say so). | List price month to month. Up to 200 minutes of generation (about $0.075/min at list); 15,000 characters per generation; 1 professional voice slot; commercial use. Annual billing was $66 at capture under a 3-months-off plus holiday 50%-off promotion ($5.50/mo effective). | Up to 1,620 minutes (about $0.062/min at list); 3 team seats; 30,000 characters per generation; 5 professional voice slots; 7-day money-back guarantee. Annual was $450 at capture ($37.50/mo effective). | Card says up to 6,250 minutes, which at $999 works out to $0.16/min and contradicts the same page's FAQ (600 to 625 credits a minute would give 40,000 minutes from 25M credits). 10 team seats; 15 professional voice slots. Annual was $8,988 at capture ($749/mo effective). Confirm the minutes in writing before buying. | Custom volume pricing billed annually: pay-as-you-go billing with organisation-level controls, zero data retention, on-premise deployment, a SOC 2 line on the card, custom SSO listed as coming soon. |
- Free Free8,000 credits · ≈ 9 min of audio
Card sizes it at up to 7 minutes of generation; 500 characters per generation; 3 public voice slots; no card needed. Voice cloning and commercial use are crossed out on the card: free output is personal use only (the FAQ and the terms both say so).
- Plus $15/mo250,000 credits · ≈ 4.6 hrs of audio
List price month to month. Up to 200 minutes of generation (about $0.075/min at list); 15,000 characters per generation; 1 professional voice slot; commercial use. Annual billing was $66 at capture under a 3-months-off plus holiday 50%-off promotion ($5.50/mo effective).
- Pro $100/mo2,000,000 credits · ≈ 37 hrs of audio
Up to 1,620 minutes (about $0.062/min at list); 3 team seats; 30,000 characters per generation; 5 professional voice slots; 7-day money-back guarantee. Annual was $450 at capture ($37.50/mo effective).
- Max $999/mo25,000,000 credits · ≈ 463 hrs of audio
Card says up to 6,250 minutes, which at $999 works out to $0.16/min and contradicts the same page's FAQ (600 to 625 credits a minute would give 40,000 minutes from 25M credits). 10 team seats; 15 professional voice slots. Annual was $8,988 at capture ($749/mo effective). Confirm the minutes in writing before buying.
- Enterprise Custom—
Custom volume pricing billed annually: pay-as-you-go billing with organisation-level controls, zero data retention, on-premise deployment, a SOC 2 line on the card, custom SSO listed as coming soon.
Credits are a prepaid balance you spend as you generate, like topping up a card. Here about one credit buys one character of speech and roughly 1,000 characters is a minute, so the bracketed figure is the speaking time each plan covers.
Prices in USD as set by the vendor · last checked 2026-07-11 · vendor pricing →
Slide your expected monthly volume to see roughly what Fish Audio would cost.
A rough estimate from Fish Audio's sourced rates, not a quote. Always confirm on the vendor's own pricing page before you commit.
At a glance
- Speech-to-text
- Fish Audio (transcribe-1)
- Text-to-speech
- Fish Audio · Bring your own voice: you can upload or clone a custom voice instead of being limited to the platform's stock ones.
- Languages
- en, zh, ja, de, fr, es, ko, ar, ru, nl, it, pl, pt
- Integrations
- REST API, WebSocket streaming TTS, Python SDK, JavaScript/TypeScript SDK, Speech-to-text API, Voice Design API, Self-host (open-source Fish Speech weights)
Compliance
Our full take
Fish Audio is a voice generator for produced audio, meaning speech you script, generate and edit before anyone hears it: voiceovers, audiobooks, character and story work. You paste text, pick a voice from a community library the homepage puts at 2,000,000-plus voices or clone your own from a claimed 15-second clip, and export the result. Behind it sits Hanabi AI Inc., a Delaware company, the OpenAudio research line, and the open-source Fish Speech project that made these models some of the most watched in speech AI. What it is not: a phone-agent platform. There are no phone numbers, no SIP trunking (plugging in your own phone carrier) and no per-minute call product anywhere on the site.
Subscriptions are priced in credits, and here are the numbers as the cards print them. Free is $0 for 8,000 credits a month, sized at up to 7 minutes of generated audio. Plus is $15 a month list for 250,000 credits (up to 200 minutes, so about $0.075 a minute). Pro is $100 for 2,000,000 credits (up to 1,620 minutes, about $0.062 a minute). Max is $999 for 25,000,000 credits, and Enterprise is custom volume pricing with zero data retention and on-premise deployment on the card. At capture the page was mid-promotion: a countdown banner priced annual billing at 3 months off plus a holiday 50% off, which is how the cards showed $5.50, $37.50 and $749 a month against struck-through list prices. We record list rates and treat the promotion as promotional.
Now the part that needs clearing up, because the page disagrees with itself. The pricing FAQ says a minute of generation costs roughly 600 to 625 credits. The Plus and Pro cards imply about 1,250 credits a minute (250,000 over 200 minutes; 2,000,000 over 1,620). The Max card implies 4,000 (25,000,000 over 6,250 minutes), which at $999 a month would make the biggest plan cost about $0.16 a minute, more than double the entry tier. No pricing ladder works that way, so something on that card is off, but we record what the page prints and flag it rather than guess. If you are sizing a plan on minutes, get the credits-per-minute figure from Fish Audio in writing first.
Developers get a much cleaner number. The API charges $15 per million UTF-8 bytes on every current model (s1, s2-pro and the recommended s2.1-pro). A byte is effectively a character if you write in English; accented and non-Latin scripts use 2 to 4 bytes per character, so they cost proportionally more. That works out at about $0.015 per 1,000 English characters, and by the docs’ own conversion (a million bytes is roughly 12 hours of speech) about $0.021 per minute of generated audio, the derived figure we carry as this page’s per-minute rate. Against the narration rates we track, that is near the bottom of the table: ElevenLabs stores at $0.10 per 1,000 characters, Cartesia at $0.035, and only Speechify is cheaper at $0.010. Two catches: pay-as-you-go API access needs a paid subscription, and concurrency is earned with spend (5 parallel requests below $100 of lifetime billing, 15 from $100, 50 from $1,000). There is also a free s2.1-pro-free API model for development with fair-use limits, speech-to-text at $0.36 per audio hour, and a Voice Design endpoint at $0.01 a request.
The models are the reason to care. The flagship is s2.1-pro, which the docs credit with 83 languages; s2-pro (80-plus languages) claims roughly 100ms time-to-first-audio on datacentre hardware; the older s1 covers 13 languages with 64-plus emotion tags you write straight into the script in brackets. Note the language claim moves depending on where you read: the homepage still says 30-plus, the docs say 83 for the flagship, and the S2 release post says the training set spans about 50. We store the docs’ flagship figure and flag the spread. The open-source angle is real, not marketing: in March 2026 Fish Audio published the S2 weights, fine-tuning code and a streaming inference engine on GitHub and Hugging Face, and a dated third-party review from June 2026 rates it the leading open-weight text-to-speech family of 2026 on blind preference, while noting ElevenLabs keeps the polish and library advantage. Those are the reviewer’s findings and the vendor’s benchmarks, not ours; we have not run scored listening tests yet.
Could you run a live agent on it? The pieces exist: websocket streaming, that 100ms first-audio claim, a conversational-chatbots solutions page. But there is no telephony, no agent product with a call rate, and the third-party review places real-world latency behind realtime specialists like Cartesia. So on this site Fish Audio registers as a narration option, not a call platform. Wire it into a voice-agent stack as the speech layer if you like the voices; buy a platform if you need the phone part handled for you.
On compliance we record formal claims only. GDPR: yes. The privacy policy carries a proper section for EEA, Swiss and UK users with the full set of data-subject rights and EU standard contractual clauses for transfers. SOC 2: the Enterprise card lists it as a plan feature, but there is no certificate, report or trust page anywhere we could find, so the box stays unticked, the same standard we hold every vendor to. HIPAA: not mentioned, so no. One rights point that catches people: free-plan output is personal use only (the card crosses out commercial use, and the FAQ and the terms agree), and commercial rights on paid plans apply to verified voices you own.
My read: for produced audio the per-character maths is genuinely hard to beat, with open weights as your exit route if the hosted service ever moves against you. Set against that, the plan page’s credit arithmetic needs cleaning up before you commit to a big tier, the compliance paperwork is thin below Enterprise, and none of the certification claims are documented. Buy it for voiceover, audiobook and character work; do not buy it expecting a phone agent. We have not placed test calls or run scored listening tests on Fish Audio, so it carries no Voxrater benchmark numbers and the scores here are an editorial preview. Everything above comes from the vendor’s plan page, docs, terms, privacy policy and blog plus one dated third-party review, captured 2026-07-11, with the pricing screenshot in evidence.
Fish Audio compared
Our in-depth pieces that put Fish Audio side by side with the field, with the sourced numbers and a clear pick.
Alternatives to Fish Audio
Other platforms that overlap with Fish Audio on the same kind of work, ranked by how many capabilities they share, then by cheaper all-in cost per minute. Compare any of them side by side on the compare page.
Tracking Fish Audio? Get the next test result
We re-test and re-price the platforms we cover. Join the list and the next dated update lands in your inbox.
Newsletter launching soon.
Sources
- Fish Audio plan page in USD (screenshot in evidence/; the page geo-priced to GBP and we switched its own currency picker to USD): Free $0 (8,000 credits, up to 7 min, commercial use crossed out), Plus $15/mo list (250,000 credits, up to 200 min), Pro $100/mo (2,000,000 credits, up to 1,620 min), Max $999/mo (25,000,000 credits, card says up to 6,250 min), Enterprise custom (zero data retention, on-premise, SOC 2 line); countdown promotion on annual billing ($66/$450/$8,988); FAQ says 600 to 625 credits per minute and free output is personal-only. · captured 2026-07-11
- API pricing docs: every current TTS model (s1, s2-pro, s2.1-pro) at $15.00 per 1M UTF-8 bytes, with s2.1-pro-free at $0 for development; '1M UTF-8 bytes is approximately 180,000 English words, or about 12 hours of speech'; speech-to-text transcribe-1 at $0.36 per audio hour; voice-design-1 at $0.01 per successful request; concurrency tiers 5/15/50 parallel requests at <$100/$100+/$1,000+ lifetime spend. · captured 2026-07-11
- Models overview docs: s2.1-pro recommended for production (83 languages), s2-pro (80+ languages, 100ms time-to-first-audio), s1 (13 languages: English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, Portuguese; 64+ emotion expressions in bracket syntax); natural-language emotion control and multi-speaker dialogue on the S2 line. · captured 2026-07-11
- Homepage: '2,000,000+ voices' community library, '30+ languages' umbrella claim (the docs claim 83 for the flagship; we flag the spread), voice cloning from a 15-second clip, product list (text-to-speech, speech-to-text, voice cloning, voice changer, story studio, audio separation, audio translation, sound effects) and an end-to-end voice-agent pitch with no priced telephony product. · captured 2026-07-11
- Affiliate terms: 20% of sale value for a year after a new customer's first purchase (recurring for 12 months on monthly subscriptions, one-time on annual), performance upgrades to 30%, monthly NET-15 payouts by FirstPromoter via PayPal or Wise, standard plans only. No attribution/cookie window stated anywhere. · captured 2026-07-11
- Terms of use (effective 2024-08-18): operator is Hanabi AI Inc., a Delaware corporation; free users are licensed for internal, personal, non-commercial use only; paid users are licensed for commercial use; no pro-rated refunds. · captured 2026-07-11
- Privacy policy (effective 2024-08-28): formal GDPR section for EEA/Swiss/UK users with the full set of data-subject rights, EU legal bases and standard contractual clauses for transfers; CCPA section; contact [email protected]. No SOC 2, HIPAA or ISO certification mentioned. · captured 2026-07-11
- Vendor blog (2026-03-09): S2 model weights, fine-tuning code and an SGLang-based streaming inference engine open-sourced on GitHub (fishaudio/fish-speech) and Hugging Face (fishaudio/s2); about 100ms time-to-first-audio on NVIDIA H200; training set 10M+ hours across roughly 50 languages (the docs' 83-language claim is the hosted flagship's). · captured 2026-07-11
- Third-party review (verified 2026-06-25): rates Fish Audio the leading open-weight TTS family of 2026 on aggregate blind preference, confirms the $15 per 1M UTF-8 bytes API rate, notes the consumer UI trails ElevenLabs, no dubbing/lip-sync tooling, and real-world latency behind realtime specialists like Cartesia. · captured 2026-07-11