Deepgram vs. Speechmatics
AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.
Both are independent speech-to-text specialists, but Speechmatics competes on language breadth, multilingual models, and flexible on-premises and on-device deployment at low per-hour batch prices, while Deepgram competes on streaming price, agent infrastructure, and a wider voice stack with standalone text-to-speech tiers.
Introduction and methodology note
Deepgram and Speechmatics are two of the longest-running independent speech-to-text vendors. Each sells batch and real-time transcription, speaker diarization, custom vocabulary, and enterprise deployment options, and each also sells text-to-speech and a voice agent product. Both give new accounts free credit and bill by the second.
They differ in emphasis. Speechmatics lists 55+ languages, multilingual models that handle mid-conversation language switching, a three-tier accuracy model (Melia 1, Standard, Enhanced), and deployment from SaaS through containers, virtual appliances, and on-device. Deepgram lists 45+ languages on Nova models, a streaming model built for agents (Flux), three standalone text-to-speech tiers, bring-your-own agent tiers, and self-hosted deployment for Enterprise.
Methodology note: Deepgram figures come from Deepgram's own published pages, linked inline. Speechmatics figures come from Speechmatics' own pricing page, which is dated July 31, 2026, and are stated as plain text. Both were checked on September 30, 2026. Where a vendor does not publish a number on the pages reviewed, this page says "not published." Accuracy claims are self-reported by each vendor, and testing on your own audio is the only reliable comparison.
TL;DR
- Choose Deepgram for lower-cost streaming at promotional rates, Flux and the Voice Agent API with published bring-your-own tiers, three standalone text-to-speech tiers, and higher default streaming concurrency (150 WebSocket connections against 50 real-time sessions).
- Choose Speechmatics for the lowest published batch rates ($0.13/hr for multilingual Melia 1, $0.24/hr for Standard), 55+ languages, and deployment as containers, virtual appliances, or on-device.
- Streaming is close: Speechmatics Real-time Standard at $0.24/hr is below Deepgram Nova-3 at $0.288/hr promotional, while Deepgram's Flux at $0.39/hr sits above it.
Quick comparison table
| Dimension | Deepgram | Speechmatics |
|---|---|---|
| Batch speech-to-text | Nova-3 $0.0043/min ($0.258/hr); multilingual $0.0052/min ($0.312/hr) | Melia 1 (multilingual) $0.13/hr; Standard $0.24/hr; Enhanced $0.40/hr |
| Streaming speech-to-text | Nova-3 $0.0048/min promotional ($0.288/hr; regular $0.0077/min, $0.462/hr); Flux English $0.0065/min ($0.39/hr) | Real-time Standard $0.24/hr; Real-time Enhanced $0.43/hr |
| Other listed STT model | Whisper Large pre-recorded $0.0048/min | Linden 1 $0.16/hr (previously $0.21/hr) |
| Languages | 45+ (Nova models) | 55+ for transcription; 69 translation pairs |
| Text-to-speech | Flux TTS $0.045/1k characters; Aura-2 $0.030/1k; Aura-1 $0.015/1k | $0.011/1k characters; English, with more languages listed as coming |
| Voice agents | Voice Agent API $0.075/min Standard; $0.050/min with BYO LLM and TTS | Voice Agent API offered; rate not published on the pricing page; 3, 6, or unlimited concurrent conversations by plan |
| Translation | Not published | Bolt-on $0.65/hr |
| Audio analysis add-ons | Summarization, topics, sentiment, and intent token-priced; entity detection $0.0017/min; redaction $0.0020/min | Summaries $0.12/hr; sentiment $0.12/hr; topics $0.20/hr; chapters $0.40/hr |
| Free entry point | $200 credit, no credit card | $100 credit, no credit card |
| Volume pricing | Growth prepaid from $4K/year, up to 20% off | 20% off above 500 hours per month per model type; further discounts from 24,000 hours per year; 33% off with opt-in model training |
| Streaming concurrency | Up to 150 WebSocket (pay-as-you-go); 225 (Growth) | 50 real-time sessions (Pro); 10 file jobs per second |
| Deployment | Cloud, EU endpoint, self-hosted (Enterprise) | SaaS in US, EU, or Australia; Enterprise private cloud, container, virtual appliance, on-device |
| Compliance | SOC 2 Type 1 and 2, HIPAA BAAs (Enterprise), GDPR, CCPA, PCI | SOC 2 Type II, ISO/IEC 27001:2022, GDPR, HIPAA |
Granular differentiator table: accuracy tiers and batch price
Speechmatics sells speech-to-text in named accuracy tiers. The table sets each tier against the closest Deepgram rate on an hourly basis.
| Speechmatics tier | Price per hour | Closest Deepgram rate | Deepgram price per hour | Lower price |
|---|---|---|---|---|
| Batch Melia 1 (multilingual) | $0.13 | Nova-3 multilingual pre-recorded | $0.312 | Speechmatics, by about 58% |
| Batch Standard | $0.24 | Nova-3 pre-recorded | $0.258 | Speechmatics, by about 7% |
| Batch Enhanced | $0.40 | Nova-3 pre-recorded | $0.258 | Deepgram, by about 35% |
| Real-time Standard | $0.24 | Nova-3 streaming | $0.288 promotional; $0.462 regular | Speechmatics at promotional; Speechmatics by about 48% at regular |
| Real-time Enhanced | $0.43 | Nova-3 streaming | $0.288 promotional; $0.462 regular | Deepgram at promotional; Speechmatics by about 7% at regular |
Speechmatics is the cheaper option on every batch tier except Enhanced, and on real-time Standard. Deepgram is cheaper where Enhanced accuracy is required at the promotional streaming rate, and on batch when Speechmatics Enhanced is the comparison. Deepgram's pre-recorded rate already includes diarization and smart formatting, and Speechmatics lists speaker diarization among its speech-to-text features without a separate charge. Because the two vendors define accuracy tiers differently, compare word error rate on your own audio rather than matching tier names.
Why teams choose Deepgram over Speechmatics
Agent-oriented recognition with published agent pricing
Flux adds model-integrated end-of-turn detection for conversational agents (streaming feature overview). The Voice Agent API publishes per-minute tiers from $0.075 Standard to $0.050 with bring-your-own LLM and TTS. Speechmatics offers a Voice Agent API and lists conversation concurrency, but does not publish an agent rate on its pricing page.
Standalone text-to-speech in three tiers
Deepgram sells Aura-1, Aura-2, and Flux TTS at $0.015 to $0.045 per 1k characters. Flux TTS runs a real-time WebSocket for live agents, a REST transport for fixed audio, and interruption reporting that states what the user heard (Flux TTS overview). Speechmatics lists one text-to-speech rate at $0.011/1k characters, currently for English.
Higher default streaming concurrency
Deepgram publishes up to 150 WebSocket connections on pay-as-you-go and 225 on Growth for speech-to-text (getting started guide). Speechmatics Pro allows 50 concurrent real-time sessions, and unlimited sessions require an Enterprise agreement.
If a voice agent stack with published tiers fits your roadmap, start a free trial with $200 in Deepgram credits.
Best use cases for Deepgram
Real-time voice agents
Flux and the Voice Agent API cover turn detection, orchestration, and synthesis with tiers that let teams supply their own LLM and voice.
How Speechmatics compares: Its Voice Agent API and Linden 1 model are listed, with conversation concurrency of 3 (Free), 6 (Pro), or unlimited (Enterprise), and no published per-minute agent rate.
High-concurrency streaming products
Live captions, agent assist, and call analytics benefit from 150 to 225 concurrent WebSocket connections without an Enterprise contract.
How Speechmatics compares: Pro includes 50 real-time sessions, and higher concurrency requires sales.
Voice products needing recognition and synthesis in one account
STT, three TTS tiers, and agents share one console, one bill, and one SDK family.
How Speechmatics compares: It sells recognition, English text-to-speech, and voice agents as well, at a lower text-to-speech rate.
Self-hosted Kubernetes deployment
Self-hosted STT, TTS, and Voice Agent deployment is published for Enterprise customers.
How Speechmatics compares: It lists private cloud, container, virtual appliance, and on-device options, which is a wider menu for edge and embedded scenarios.
Best use cases for Speechmatics
Multilingual and code-switching transcription
Melia 1 transcribes multilingual speech in a single transcript without selecting a language, including speakers who switch mid-conversation, at $0.13/hr for batch. With 55+ languages and 69 translation pairs, it fits global call archives and media. Deepgram's Nova models list 45+ languages.
On-device, private, and offline deployment
Speechmatics lists deployment from SaaS in three regions to containers, virtual appliances, and on-device models, with on-premises text-to-speech on Enterprise. Organizations that must process audio entirely inside their own environment, or on the device, have a wider published menu here.
Platform overview of Deepgram
Speech-to-text with Nova-3
Nova-3 serves pre-recorded and streaming audio in 45+ languages with diarization, keyterm prompting, and language detection. See the models and languages overview.
Flux conversational recognition
Flux is a streaming model with built-in end-of-turn detection for agents, in English and multilingual variants. Start with the Flux quickstart.
Text-to-speech with Aura-2 and Flux TTS
Synthesis is sold standalone at $0.015 to $0.045 per 1k characters. The text-to-speech guide covers requests and voices.
Voice Agent API
One WebSocket handles listening, reasoning, and speaking, with bundled and bring-your-own tiers. Read the Voice Agent documentation.
Audio Intelligence
Summarization, topic detection, sentiment, and intent recognition are token-priced, and entity detection and redaction are billed per minute. See the Audio Intelligence product page.
Integrations
Deepgram documents LiveKit, Pipecat, Twilio, Amazon Connect, and Genesys integrations. The LiveKit integration guide shows a working agent pattern.
How to choose
| If your priority is | Lean toward | Because |
|---|---|---|
| Lowest multilingual batch price | Speechmatics | Melia 1 at $0.13/hr |
| Voice agents with published per-minute tiers | Deepgram | $0.075 to $0.050/min against no published rate |
| Code-switching within a call | Speechmatics | Melia 1 handles mid-conversation switching |
| Higher default streaming concurrency | Deepgram | 150 to 225 WebSocket connections against 50 sessions |
| On-device or virtual appliance deployment | Speechmatics | Published deployment menu |
| Three standalone TTS tiers | Deepgram | Aura-1, Aura-2, and Flux TTS |
| Lowest English TTS rate | Speechmatics | $0.011/1k characters against $0.015 to $0.045 |
| Languages beyond 45 | Speechmatics | 55+ languages |
Switching from Speechmatics to Deepgram
- Set up the account. Create a Deepgram project and key. The Deepgram Python SDK is the quickest route to a first request, and the $200 credit covers a parallel test.
- Check your languages first. List every language you transcribe and confirm each appears in Deepgram's language documentation. Speechmatics covers 55+, and any language outside Nova's 45+ should stay where it is.
- Move batch jobs. Speechmatics batch jobs map to a single POST to the pre-recorded endpoint with
diarize=trueandsmart_format=true. Custom dictionary entries map to keyterm prompting. - Move real-time sessions. Speechmatics real-time sessions map to a WebSocket connection to Nova-3, or to Flux for agents. Standard and Enhanced tier choices become model choices, so retest accuracy rather than assuming a mapping.
- Rework add-ons. Speechmatics translation, chapters, and summary bolt-ons have different shapes than Deepgram's token-priced intelligence features, and translation has no published Deepgram equivalent.
- Run in parallel. Mirror one to two weeks of traffic, and compare word error rate on your own audio, latency at your percentiles, and invoice line items. A partial migration is a legitimate end state, such as moving agents and leaving multilingual archives on Speechmatics.
FAQ
Is Deepgram cheaper than Speechmatics?
It depends on the mode. Speechmatics is cheaper on batch ($0.13/hr for Melia 1 and $0.24/hr for Standard against $0.258/hr for Nova-3) and on Real-time Standard ($0.24/hr). Deepgram is cheaper on streaming at the promotional Nova-3 rate against Real-time Enhanced ($0.288/hr against $0.43/hr).
Which is more accurate?
Both vendors publish their own accuracy claims and define tiers differently. Test both on a sample of your own audio and accents.
How do the free tiers compare?
Deepgram gives a $200 credit and Speechmatics gives a $100 credit, both without a credit card.
Does Speechmatics offer text-to-speech and voice agents?
Yes. It lists text-to-speech at $0.011/1k characters for English and a Voice Agent API, with no per-minute agent rate on its pricing page.
Can I self-host either platform?
Both offer it on Enterprise terms. Deepgram publishes self-hosted STT, TTS, and Voice Agent deployment, and Speechmatics lists private cloud, container, virtual appliance, and on-device options.
How hard is migration?
Transcription moves in days because both are REST plus WebSocket APIs with SDKs. Language coverage checks and accuracy tier retesting take the most time.
Both vendors give free credit, so the fastest way to settle this comparison is a test on your own audio and languages. When you are ready to size it for production volume, talk to the Deepgram sales team about Growth and Enterprise plans.