Speechmatics
Enterprise speech-to-text with very broad language coverage and real on-prem options, for teams who self-host.
Paid link, we may earn a commission. How this works.
Scored on the same voice-agent rubric as the full platforms, so a building block like this scores low on the axes it does not address. Read its value score against its job.
See how it stacks up · Full rankings →The languages-and-deployment specialist. Speechmatics turns speech into text in 56+ languages and will run inside your own data centre, not just its cloud. It is one building block though, not a whole phone agent. No voice, no language model, no phone line.
About $0.00 to 0.01 for a minute of conversation, once the phone line and the AI are added in.
That's roughly $0.24–0.43 an hour. Plans: $0/mo (Free).
Pricing
Show the cost breakdown
| What the platform charges to run the agent, before the phone line and the AI usage are added on. | — |
|---|---|
| The step that turns what the caller says out loud into text the AI can read. | $0.00 /min |
| The AI 'brain' that reads what the caller said and works out what to say back. | — |
| The step that turns the AI's written reply back into a spoken voice. | — |
| The phone line itself: the service that connects the call to a real phone number. Usually billed on top of the platform. | — |
| The total you actually pay for one minute of conversation once every piece is added up: the platform, the AI, the voice and the phone line. | $0.00–0.01 /min |
Speechmatics prices per HOUR of audio, not per minute. The 2026-07-11 rate card (read from the dated capture; the live table is JavaScript-rendered) gives real-time STT at $0.24/hr Standard and $0.43/hr Enhanced, batch at $0.24/hr Standard and $0.40/hr Enhanced, so the per-minute band here is $0.004 (Standard) to about $0.0072 (Enhanced). A new multilingual model, Melia 1 (launched 2026-06-17), is a batch-only production preview advertised as low as $0.129/hr; that rate is tied to the model-training data-sharing discount (the rate table renders with the 33% discount toggle on by default), so the undiscounted rate is higher. Speechmatics also ships text-to-speech at $0.011 per 1,000 characters (English at launch). The free tier is now 3,000 minutes a month (50 hours), split 1,200 real-time plus 1,800 batch, plus about 1M free TTS characters, no card needed. A 20% volume discount applies above 500 hours a month per service. This is primarily speech-to-text: there is no language model and no telephony, so those components are 0 here. To run a full phone agent you add an LLM and a phone line separately, each a cost on top.
Every plan in one place: the monthly fee, what each one includes, and the features it unlocks. Anything beyond a plan's allowance, or on a pay-as-you-go tier, is billed at the per-minute rate above. A blank in the features means the vendor's plan page does not state it for that plan, not that it is unavailable.
| Free | Pro | Enterprise | |
|---|---|---|---|
| Price | Free | — | Custom |
| Included | 3,000 minutes | Pay per use | — |
| Plan notes | 3,000 free minutes (50 hours) per month, split 1,200 real-time + 1,800 batch, no card required, plus ~1M free TTS characters; 2 concurrent real-time sessions | Pay-as-you-go on usage. Real-time STT $0.24/hr Standard, $0.43/hr Enhanced; batch $0.24/hr Standard, $0.40/hr Enhanced; TTS $0.011 per 1,000 characters. New Melia 1 multilingual model in batch-only production preview, advertised as low as $0.129/hr (a rate tied to the model-training data-sharing discount; the undiscounted rate is higher). Capped at 6,000 hours/month, 50 concurrent real-time sessions. | Custom pricing, no rate limits, on-prem/container deployment, volume discounts from 24,000 hours/year |
| What each plan unlocks | |||
| API access | Yes | Yes | — |
| Concurrent calls | 2 real-time sessions | 50 real-time sessions | — |
| Priority support | — | — | Custom deployment + volume pricing |
- Free Free3,000 minutes
3,000 free minutes (50 hours) per month, split 1,200 real-time + 1,800 batch, no card required, plus ~1M free TTS characters; 2 concurrent real-time sessions
- API access
- Yes
- Concurrent calls
- 2 real-time sessions
- Priority support
- —
- Pro —Pay per use
Pay-as-you-go on usage. Real-time STT $0.24/hr Standard, $0.43/hr Enhanced; batch $0.24/hr Standard, $0.40/hr Enhanced; TTS $0.011 per 1,000 characters. New Melia 1 multilingual model in batch-only production preview, advertised as low as $0.129/hr (a rate tied to the model-training data-sharing discount; the undiscounted rate is higher). Capped at 6,000 hours/month, 50 concurrent real-time sessions.
- API access
- Yes
- Concurrent calls
- 50 real-time sessions
- Priority support
- —
- Enterprise Custom—
Custom pricing, no rate limits, on-prem/container deployment, volume discounts from 24,000 hours/year
- API access
- —
- Concurrent calls
- —
- Priority support
- Custom deployment + volume pricing
Each plan bundles a set amount of talk time a month.
Prices in USD as set by the vendor · last checked 2026-07-11 · vendor pricing →
At a glance
- Speech-to-text
- Speechmatics (Standard / Enhanced), Speechmatics Melia 1 (multilingual, batch-only preview)
- Text-to-speech
- Speechmatics TTS
- Languages
- en, es, fr, de, it, pt, nl, pl, ru, ar, hi, zh, ja, ko, cy
- Integrations
- Real-time API (streaming), Batch API (recorded files), On-prem containers (CPU / GPU), Kubernetes self-host, Virtual Appliance (on-prem VM), Native SDKs
Compliance
Our full take
Speechmatics is a speech-to-text engine first, and that is the whole point to get straight. It listens to audio and writes down the words. It now also offers its own text-to-speech, but it does not generate a reply (there is no language model) and it does not dial a phone. So if you are shopping for a finished voice agent that answers your calls, this is not that. It is one of the parts you would build that agent from, and it is a good one.
Where it earns its place is languages. Speechmatics transcribes 56+ languages off a single model, which means you get the regional accents and dialects (Brazilian Portuguese, Canadian French, and so on) without bolting on a separate pack for each. Most of the cheaper speech-to-text engines top out around seven or ten languages. Deepgram, the closest building-block vendor we cover, lists seven. If your callers speak Tagalog, Welsh, Swahili or Urdu, that gap is the entire reason to look here.
The second reason is where it runs. Most speech-to-text APIs only run in the vendor’s cloud, you send them audio and they send back text. Speechmatics will also run inside your own data centre, as a container on your own hardware (CPU or GPU), on Kubernetes, or as a pre-built virtual machine they call a Virtual Appliance. For a hospital or a bank that cannot let call audio leave the building, that on-premises option (meaning it runs on your own servers, not someone else’s cloud) is often a hard requirement, not a nice-to-have. It is the kind of thing you cannot retrofit, so it matters that it is there from the start.
Now the pricing, and here is the bit that trips people up. Speechmatics bills per hour of audio, not per minute like the agent platforms. The pricing page now headlines the Pro plan at “from $0.129 an hour”, but read that carefully: it is the new Melia 1 batch rate with a data-sharing discount already applied (more on both below). The dependable floor for the established models is $0.24 an hour on Standard, which works out at about $0.004 a minute, the figure shown at the top of this page. Treat that as the floor, not the average. The real rate climbs with the model you pick and the mode you run.
Here is roughly how it splits, from the current rate card. Real-time transcription (live, as the audio streams in) is $0.24 an hour on the Standard model and $0.43 on the higher-accuracy Enhanced model. Batch transcription (a recorded file after the fact) is $0.24 Standard and $0.40 Enhanced. In per-minute terms that is about $0.004 to $0.0072 depending on model and mode. The rate table loads through JavaScript, so we read these figures from a dated capture rather than a plain page fetch. The free tier has grown to 3,000 minutes a month (50 hours), split 1,200 real-time and 1,800 batch, no card needed, plus around a million free text-to-speech characters, and a 20% volume discount once you cross 500 hours a month.
The new model is Melia 1, launched 2026-06-17. It transcribes multilingual audio in one pass, including speakers who switch language mid-sentence, across 55+ languages by Speechmatics’ own description, with no need to pick a language up front. Two caveats before you build on it. First, it is a batch-only production preview for now (real-time is promised, not shipped), so it cannot power a live phone agent yet. Second, the advertised as low as $0.129 an hour is tied to the model-training toggle, a setting that shares your audio data with Speechmatics in exchange for a discount; leave that off and the rate is higher. Speechmatics’ launch post also claims Melia beats Deepgram, Microsoft and AssemblyAI on the FLEURS multilingual test set. That is the vendor marking its own homework, so treat it as a claim, not a result.
One honest caveat on cost. That per-minute number looks tiny next to a $0.06-a-minute agent platform, and it is, but it is not comparing like for like. Speechmatics is charging you for one job, the transcription. The platforms are charging for transcription plus the language model plus the voice plus the phone line bundled together. To build a full phone agent on Speechmatics you still have to pay for an LLM and telephony separately, though it now has its own text-to-speech ($0.011 per 1,000 characters) so the voice no longer has to come from elsewhere. Add those up and the real per-minute cost lands a lot closer to the bundled platforms than the $0.004 headline suggests.
On compliance, Speechmatics is unusually well-documented for a building block. Its own security page states SOC 2 Type II, ISO/IEC 27001:2022, GDPR and full HIPAA compliance, with AES 256 encryption at rest and TLS 1.2 or higher in transit, plus a public trust centre where you can pull the actual reports. We have ticked HIPAA, SOC 2 Type II and GDPR here because the vendor states them directly. We left SOC 2 Type I unticked: the page names Type II, not Type I, and we do not assume one from the other. For a regulated buyer, that combination of on-prem deployment plus written certifications is the strong card.
My read: Speechmatics is the one you reach for when language coverage or on-premises deployment is non-negotiable, and you have the engineering to assemble the rest of the agent around it. The voice-quality and ease-of-use scores sit lower here than for a finished platform, and that is fair, this is infrastructure, not a product you switch on. If you just want calls answered without standing up your own stack, a bundled platform will get you there faster. If you need to transcribe twenty languages, or keep the audio on your own servers, very little else competes.
The 1 to 10 scores on this page are an editorial preview, our provisional read to get the framework in place, not a measured result. We have not run Speechmatics through our own test calls yet, so there is no Voxrater latency figure here. The pricing, language, deployment and compliance detail is sourced from Speechmatics’ own pricing, security, languages and deployments pages, first captured 2026-05-31 and most recently re-captured 2026-07-11.
Alternatives to Speechmatics
Other platforms that overlap with Speechmatics on the same kind of work, ranked by how many capabilities they share, then by cheaper all-in cost per minute. Compare any of them side by side on the compare page.
Tracking Speechmatics? Get the next test result
We re-test and re-price the platforms we cover. Join the list and the next dated update lands in your inbox.
Newsletter launching soon.
Sources
- Speechmatics pricing re-captured 2026-07-11 (JS-rendered table read from the dated screenshot; the table renders with the 'Model Training: enable for 33% discount' toggle ON by default, so the displayed per-hour figures are the discounted ones): free tier grown to 3,000 STT min/mo (50 hrs), split 1,200 real-time + 1,800 batch, plus ~1M TTS chars; 56+ languages; Batch Melia 1 listed at $0.129/hr alongside Standard/Enhanced. · captured 2026-07-11
- Melia 1 launch post (published 2026-06-17): multilingual model with mid-utterance code-switching across 55+ languages, batch-only production preview ('real-time on the way'), advertised as low as $0.129/hr with 10 free hours a month; the FLEURS accuracy wins over Deepgram, Microsoft and AssemblyAI are Speechmatics' own claims, not independent results. · captured 2026-07-11
- Startup Program page, checked 2026-07-11: the old become-a-partner URL now redirects here; up to $50k in usage credits, cohorts capped at 20 startups, under $10M raised; no partner, reseller or affiliate content remains. · captured 2026-07-11
- Speechmatics pricing re-captured 2026-06-15 (rate table is JS-rendered; figures read from the dated screenshot): real-time STT $0.24/hr Standard, $0.43/hr Enhanced; batch $0.24/$0.40; free tier now 2,400 min/mo; models named Standard/Enhanced. · captured 2026-06-15
- Speechmatics text-to-speech (new): $0.011 per 1,000 characters, English at launch. · captured 2026-06-15
- Speechmatics pricing verified 2026-06-02: Pro from $0.24/hr (= $0.004/min), 2,400 free minutes/mo; speech-to-text only, so no per-minute voice-output rate. · captured 2026-06-02
- Speechmatics pricing page re-captured 2026-06-02 for the quarterly re-verification (screenshot in evidence/). · captured 2026-06-02
- Speechmatics pricing page: per-plan features (Free, Pro, Enterprise), Pro from $0.24/hr, Free 480 min/month, 6,000 hr/month cap, 24,000 hr/year enterprise discount · captured 2026-05-31
- Speechmatics security page: SOC 2 Type II, ISO/IEC 27001:2022, GDPR and HIPAA claims, AES 256 / TLS 1.2+, Azure + on-prem · captured 2026-05-31
- Speechmatics languages page: 55+ languages for speech-to-text with accent/dialect coverage · captured 2026-05-31
- Features and deployments: SaaS, on-prem containers (CPU/GPU), Kubernetes, Virtual Appliance, real-time + batch · captured 2026-05-31
- Speechmatics partner programme: build/market/sell tracks plus partner marketplace · captured 2026-05-31