Deepgram vs. AssemblyAI
AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.
Both are strong speech-to-text APIs. The practical difference for most teams is scope: Deepgram sells a full voice stack (STT, standalone TTS, composable voice agents, and self-hosted deployment), while AssemblyAI pairs low-cost transcription with a deep menu of audio-understanding add-ons.
Introduction and methodology note
Deepgram and AssemblyAI are the two most commonly shortlisted independent speech-to-text APIs. Both offer batch and real-time transcription, speaker diarization, redaction, and language coverage well beyond English, and both are priced per unit of audio with usable free tiers. If your only requirement is "turn audio files into text at a fair price," you can build a working product on either.
The differences show up at the edges of that requirement. Deepgram's catalog extends into standalone text-to-speech (Aura-2 and Flux TTS), a Voice Agent API with bring-your-own-model tiers, and self-hosted deployment for teams that cannot send audio to a vendor cloud. AssemblyAI extends in a different direction: aggressively priced async transcription and a large a-la-carte speech understanding menu (sentiment, topics, content moderation, PII handling) billed per feature per hour. Which edge matters depends on what you are building.
Methodology note: Deepgram figures come from Deepgram's own published pricing and developer documentation, and AssemblyAI figures come from AssemblyAI's own published pricing and product pages, stated as plain text. Every price and feature claim was checked in September 2026. Where a vendor does not publish a number, this page says "not published." Accuracy claims are self-reported by each vendor, and testing on your own audio is the only reliable comparison.
TL;DR
- Choose Deepgram if you need the full voice loop (STT plus TTS plus agents), multilingual real-time streaming, composable voice-agent pricing, or self-hosted deployment. Streaming Nova-3 is currently $0.0048/min (promotional; regular $0.0077/min), and pre-recorded is $0.0043/min with diarization and smart formatting included.
- Choose AssemblyAI if your workload is primarily async transcription plus audio analysis. Its published base rates are $0.15/hr (Universal-2) to $0.21/hr (Universal-3.5 Pro) for batch and $0.15/hr for Universal-Streaming, with understanding features added per hour as needed.
- Voice agents cost the same headline rate on both ($4.50/hr bundled), but Deepgram publishes BYO-LLM and BYO-TTS tiers down to $0.050/min, while AssemblyAI publishes one bundled rate.
Quick comparison table
| Dimension | Deepgram | AssemblyAI |
|---|---|---|
| Real-time streaming STT | Nova-3 monolingual $0.0048/min (promotional; regular $0.0077/min); Nova-3 multilingual $0.0058/min (promotional; regular $0.0092/min); Flux English $0.0065/min (current; regular $0.0077/min) | Universal-Streaming $0.15/hr ($0.0025/min); Universal-3.5 Pro Realtime $0.45/hr ($0.0075/min) |
| Pre-recorded (batch) STT | Nova-3 monolingual $0.0043/min ($0.258/hr); multilingual $0.0052/min | Universal-2 $0.15/hr; Universal-3.5 Pro $0.21/hr |
| Standalone text-to-speech | Aura-2 $0.030/1k characters; Flux TTS $0.045/1k characters | Not sold standalone; TTS bundled inside the Voice Agent API |
| Voice agent API | $0.075/min Standard; BYO tiers from $0.050/min | $4.50/hr flat, bundled STT, built-in LLM, and TTS; no BYO tiers published |
| Speaker diarization | Included on pre-recorded; $0.0020/min streaming | +$0.02/hr async; +$0.12/hr streaming |
| Free tier | $200 credits (roughly 775 hours of Nova-3 batch at list price), no credit card | 185 hours pre-recorded plus 333 hours streaming (published free tier) |
| Self-hosted deployment | Published (Enterprise) | Not published |
| Concurrency | Up to 50 REST and 150 WSS on pay-as-you-go; 225 WSS on Growth; higher via Enterprise | Advertises unlimited concurrent streams |
| Audio understanding | Summarization, topics, sentiment, and intents (token-priced); entity detection and redaction per minute | Larger a-la-carte menu per hour: sentiment, topics, moderation, PII text and audio redaction, key phrases, translation, more |
| Compliance | SOC 2 Type 1 and 2, HIPAA BAAs (Enterprise), GDPR with EU endpoint, CCPA, PCI | Publishes SOC 2 and GDPR compliance |
Granular table: batch transcription plus add-ons (AssemblyAI's differentiator)
Like-for-like hourly cost for a common production configuration: batch transcription with speaker labels, formatting, PII redaction, and domain vocabulary.
| Line item | Deepgram (Nova-3 mono, batch) | AssemblyAI (published rates) |
|---|---|---|
| Base transcription | $0.258/hr ($0.0043/min) | $0.15/hr (Universal-2) or $0.21/hr (Universal-3.5 Pro) |
| Speaker diarization | Included | +$0.02/hr |
| Smart or custom formatting | Included | +$0.03/hr |
| PII redaction (text) | +$0.12/hr ($0.0020/min) | +$0.08/hr |
| Domain vocabulary (keyterms) | +$0.078/hr ($0.0013/min) | +$0.05/hr |
| Configured total | $0.456/hr | $0.33/hr to $0.39/hr depending on model |
AssemblyAI stays cheaper for this batch configuration even fully loaded. The picture inverts in real-time, where Deepgram's current Nova-3 streaming rate ($0.288/hr promotional) sits well under AssemblyAI's Pro-tier realtime ($0.45/hr). Price the mode you will actually run.
Why teams choose Deepgram over AssemblyAI
One vendor for the whole voice loop
Deepgram sells STT, standalone TTS (Aura-2 and Flux TTS), and a Voice Agent API from one console and one bill (Voice Agent documentation). AssemblyAI does not sell standalone TTS, so building a talking product on it means adding a second voice vendor and stitching latency budgets across providers.
Composable voice agents, not a fixed bundle
Both vendors' bundled agents cost $4.50/hr. Deepgram additionally publishes BYO tiers: bring your own TTS ($0.065/min), your own LLM, or both ($0.050/min). Teams with an existing model stack keep it and pay less, and teams without one take the bundle.
Deployment control
Deepgram publishes a self-hosted deployment path for running models inside your own infrastructure (self-hosted Voice Agent guide), plus a dedicated EU endpoint for GDPR data residency. AssemblyAI does not publish a self-hosted option, which can be disqualifying in regulated or air-gapped environments.
If one vendor for the full voice loop fits your roadmap, start free with $200 in Deepgram credits and run your own audio through Nova-3, Flux, and the Voice Agent API before you commit.
Best use cases for Deepgram
Real-time voice agents
Flux is priced and built specifically for conversational turn-taking (Flux quickstart), and the Voice Agent API handles interruptions and barge-in over a single WebSocket.
How AssemblyAI compares: its Voice Agent API covers the same bundled loop at the same $4.50/hr, but with a single fixed configuration and no standalone TTS to reuse elsewhere.
Multilingual streaming products
Nova-3 multilingual streams at $0.0058/min promotional, and Nova models cover 45+ languages with automatic language detection (models and languages overview).
How AssemblyAI compares: it publishes multilingual streaming at its base $0.15/hr tier, but no Pro-tier multilingual realtime rate is published.
Regulated and on-prem workloads
Self-hosted deployment, HIPAA BAAs, PCI compliance, and an EU data residency endpoint are all published options.
How AssemblyAI compares: cloud API only, per its published materials, with no self-hosted path published.
High-accuracy batch with domain vocabulary
Keyterm prompting boosts recognition of product names and jargon at $0.0013/min, with diarization and smart formatting included in the base batch rate.
How AssemblyAI compares: keyterms prompting is published at +$0.05/hr and diarization at +$0.02/hr on top of a cheaper base rate, and fully configured it still totals less per hour for batch.
Best use cases for AssemblyAI
High-volume async transcription on a budget
At published rates of $0.15/hr to $0.21/hr for batch work, AssemblyAI undercuts Deepgram's $0.258/hr base for pure transcription jobs where real-time delivery does not matter. If you are transcribing podcast archives or recorded meetings at scale, this is a real cost advantage.
Audio analysis pipelines
AssemblyAI's a-la-carte understanding menu (sentiment, topics, content moderation, key phrases, translation, PII audio redaction) is broader than Deepgram's published intelligence lineup, and its LLM-over-audio tooling is a first-class product. Teams whose product is the analysis, not the transcript, may find the richer menu decisive.
Platform overview of Deepgram
Speech-to-text with Nova-3
Nova-3 serves batch and streaming in 45+ languages with diarization, smart formatting, keyterm prompting, and language detection. See the streaming feature overview.
Flux conversational recognition
Flux is streaming recognition built for voice agents, with turn-taking and conversational latency, from $0.0065/min. The getting started guide compares the streaming paths.
Text-to-speech with Aura-2 and Flux TTS
Aura-2 ($0.030/1k characters) and Flux TTS ($0.045/1k characters) are sold standalone. The Flux TTS overview covers the real-time and REST transports.
Voice Agent API
Bundled or bring-your-own tiers from $0.050/min, with interruptions and turn-taking over one WebSocket. Tiers run Standard, Custom, and Advanced, each with BYO-TTS variants.
Audio Intelligence
Summarization, topic detection, sentiment, and intent recognition, with entity detection and redaction billed per minute. See the Audio Intelligence product page.
Deployment and trust
Cloud, a dedicated EU endpoint, and self-hosted deployment, with SOC 2 Type 1 and 2, HIPAA BAAs, GDPR, CCPA, and PCI. The Flux self-hosted guide covers the GPU requirements.
How to choose
| If your priority is | Lean toward | Because |
|---|---|---|
| Real-time voice products (agents, live captions, IVR) | Deepgram | Streaming rates currently below AssemblyAI's Pro realtime tier, plus Flux and a composable agent stack |
| Cheapest possible batch transcription | AssemblyAI | Published base rates of $0.15/hr to $0.21/hr beat Deepgram's $0.258/hr, even with add-ons stacked |
| Text-to-speech or the full voice loop | Deepgram | Standalone TTS exists, while AssemblyAI's TTS is only inside its agent bundle |
| Audio analysis depth | AssemblyAI | Broader published understanding menu |
| Self-hosting, data residency, regulated industries | Deepgram | Published self-hosted path, EU endpoint, HIPAA BAAs, PCI |
| Massive concurrent streaming with no caps | Test both | AssemblyAI advertises unlimited streams, and Deepgram publishes caps that rise by plan and are negotiable at Enterprise |
Switching from AssemblyAI to Deepgram
- Map the endpoints. Batch jobs move from AssemblyAI's async transcript flow to a single POST to Deepgram's pre-recorded endpoint, and streaming moves WebSocket to WebSocket. SDKs are available for Python, JavaScript, Go, and .NET.
- Translate feature flags. Speaker labels become
diarize=true(included on batch). Formatting becomessmart_format=true(included). Custom vocabulary maps to keyterm prompting. PII redaction maps to the redact parameters. - Move webhooks. AssemblyAI's completion webhooks map to Deepgram's callback URLs on async requests. Payload shapes differ, so plan a thin adapter.
- Run in parallel. Mirror one to two weeks of production traffic to both APIs and compare word error rate on your own audio, latency at your percentiles, and invoice line items. The $200 credit typically covers this test at no cost.
- Cut over by workload. Teams commonly move streaming first, where the price and latency gap is largest, and batch last, or keep batch on the incumbent if the economics favor it. A partial migration is a legitimate end state.
FAQ
Is Deepgram or AssemblyAI cheaper for real-time streaming?
It depends on the model tier. AssemblyAI's Universal-Streaming is published at $0.15/hr ($0.0025/min), lower than Deepgram's Nova-3 monolingual streaming at $0.0048/min (promotional; regular $0.0077/min). AssemblyAI's higher-accuracy Universal-3.5 Pro Realtime is published at $0.45/hr ($0.0075/min), above Deepgram's current Nova-3 rate. Compare the tiers you would actually run, and test accuracy on your own audio before deciding on price alone.
Which is more accurate, Deepgram Nova-3 or AssemblyAI Universal?
Both vendors publish self-reported benchmarks favoring their own models, and independent results vary by audio domain, noise, and language. The only reliable answer for your workload is a side-by-side test on your own audio, and both platforms make that free to run.
Does AssemblyAI offer text-to-speech?
Not as a standalone API. Its published TTS capability is bundled inside its $4.50/hr Voice Agent API. Deepgram sells standalone TTS (Aura-2 at $0.030/1k characters, Flux TTS at $0.045/1k characters) alongside STT and agents.
Can I self-host either platform?
Deepgram publishes self-hosted deployment documentation for Enterprise customers. AssemblyAI does not publish a self-hosted or on-premises option, and its products run as a cloud API.
How hard is it to migrate from AssemblyAI to Deepgram?
Both are REST plus WebSocket APIs with official SDKs, so migration is mostly endpoint and parameter mapping, plus a thin webhook adapter. Most teams mirror traffic to both for one to two weeks before cutting over.
Which should I choose for a voice agent?
Both bundle a voice agent at $4.50/hr. Deepgram additionally publishes BYO tiers down to $0.050/min with your own LLM and TTS, and AssemblyAI publishes one bundled rate. If you want control over the LLM or voices, Deepgram's tiering is built for that.
Both platforms offer free tiers, so the fastest way to settle this comparison is a side-by-side test on your own audio. When you are ready to size it for production volume, talk to the Deepgram sales team about Growth and Enterprise plans.