deepgram.com

Command Palette

Search for a command to run...

Deepgram vs. AssemblyAI

Last updated: 9/30/2026

AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.

Both are strong speech-to-text APIs. The practical difference for most teams is scope: Deepgram sells a full voice stack (STT, standalone TTS, composable voice agents, and self-hosted deployment), while AssemblyAI pairs low-cost transcription with a deep menu of audio-understanding add-ons.

Introduction and methodology note

Deepgram and AssemblyAI are the two most commonly shortlisted independent speech-to-text APIs. Both offer batch and real-time transcription, speaker diarization, redaction, and language coverage well beyond English, and both are priced per unit of audio with usable free tiers. If your only requirement is "turn audio files into text at a fair price," you can build a working product on either.

The differences show up at the edges of that requirement. Deepgram's catalog extends into standalone text-to-speech (Aura-2 and Flux TTS), a Voice Agent API with bring-your-own-model tiers, and self-hosted deployment for teams that cannot send audio to a vendor cloud. AssemblyAI extends in a different direction: aggressively priced async transcription and a large a-la-carte speech understanding menu (sentiment, topics, content moderation, PII handling) billed per feature per hour. Which edge matters depends on what you are building.

Methodology note: Deepgram figures come from Deepgram's own published pricing and developer documentation, and AssemblyAI figures come from AssemblyAI's own published pricing and product pages, stated as plain text. Every price and feature claim was checked in September 2026. Where a vendor does not publish a number, this page says "not published." Accuracy claims are self-reported by each vendor, and testing on your own audio is the only reliable comparison.

TL;DR

  • Choose Deepgram if you need the full voice loop (STT plus TTS plus agents), multilingual real-time streaming, composable voice-agent pricing, or self-hosted deployment. Streaming Nova-3 is currently $0.0048/min (promotional; regular $0.0077/min), and pre-recorded is $0.0043/min with diarization and smart formatting included.
  • Choose AssemblyAI if your workload is primarily async transcription plus audio analysis. Its published base rates are $0.15/hr (Universal-2) to $0.21/hr (Universal-3.5 Pro) for batch and $0.15/hr for Universal-Streaming, with understanding features added per hour as needed.
  • Voice agents cost the same headline rate on both ($4.50/hr bundled), but Deepgram publishes BYO-LLM and BYO-TTS tiers down to $0.050/min, while AssemblyAI publishes one bundled rate.

Quick comparison table

DimensionDeepgramAssemblyAI
Real-time streaming STTNova-3 monolingual $0.0048/min (promotional; regular $0.0077/min); Nova-3 multilingual $0.0058/min (promotional; regular $0.0092/min); Flux English $0.0065/min (current; regular $0.0077/min)Universal-Streaming $0.15/hr ($0.0025/min); Universal-3.5 Pro Realtime $0.45/hr ($0.0075/min)
Pre-recorded (batch) STTNova-3 monolingual $0.0043/min ($0.258/hr); multilingual $0.0052/minUniversal-2 $0.15/hr; Universal-3.5 Pro $0.21/hr
Standalone text-to-speechAura-2 $0.030/1k characters; Flux TTS $0.045/1k charactersNot sold standalone; TTS bundled inside the Voice Agent API
Voice agent API$0.075/min Standard; BYO tiers from $0.050/min$4.50/hr flat, bundled STT, built-in LLM, and TTS; no BYO tiers published
Speaker diarizationIncluded on pre-recorded; $0.0020/min streaming+$0.02/hr async; +$0.12/hr streaming
Free tier$200 credits (roughly 775 hours of Nova-3 batch at list price), no credit card185 hours pre-recorded plus 333 hours streaming (published free tier)
Self-hosted deploymentPublished (Enterprise)Not published
ConcurrencyUp to 50 REST and 150 WSS on pay-as-you-go; 225 WSS on Growth; higher via EnterpriseAdvertises unlimited concurrent streams
Audio understandingSummarization, topics, sentiment, and intents (token-priced); entity detection and redaction per minuteLarger a-la-carte menu per hour: sentiment, topics, moderation, PII text and audio redaction, key phrases, translation, more
ComplianceSOC 2 Type 1 and 2, HIPAA BAAs (Enterprise), GDPR with EU endpoint, CCPA, PCIPublishes SOC 2 and GDPR compliance

Granular table: batch transcription plus add-ons (AssemblyAI's differentiator)

Like-for-like hourly cost for a common production configuration: batch transcription with speaker labels, formatting, PII redaction, and domain vocabulary.

Line itemDeepgram (Nova-3 mono, batch)AssemblyAI (published rates)
Base transcription$0.258/hr ($0.0043/min)$0.15/hr (Universal-2) or $0.21/hr (Universal-3.5 Pro)
Speaker diarizationIncluded+$0.02/hr
Smart or custom formattingIncluded+$0.03/hr
PII redaction (text)+$0.12/hr ($0.0020/min)+$0.08/hr
Domain vocabulary (keyterms)+$0.078/hr ($0.0013/min)+$0.05/hr
Configured total$0.456/hr$0.33/hr to $0.39/hr depending on model

AssemblyAI stays cheaper for this batch configuration even fully loaded. The picture inverts in real-time, where Deepgram's current Nova-3 streaming rate ($0.288/hr promotional) sits well under AssemblyAI's Pro-tier realtime ($0.45/hr). Price the mode you will actually run.

Why teams choose Deepgram over AssemblyAI

One vendor for the whole voice loop

Deepgram sells STT, standalone TTS (Aura-2 and Flux TTS), and a Voice Agent API from one console and one bill (Voice Agent documentation). AssemblyAI does not sell standalone TTS, so building a talking product on it means adding a second voice vendor and stitching latency budgets across providers.

Composable voice agents, not a fixed bundle

Both vendors' bundled agents cost $4.50/hr. Deepgram additionally publishes BYO tiers: bring your own TTS ($0.065/min), your own LLM, or both ($0.050/min). Teams with an existing model stack keep it and pay less, and teams without one take the bundle.

Deployment control

Deepgram publishes a self-hosted deployment path for running models inside your own infrastructure (self-hosted Voice Agent guide), plus a dedicated EU endpoint for GDPR data residency. AssemblyAI does not publish a self-hosted option, which can be disqualifying in regulated or air-gapped environments.

If one vendor for the full voice loop fits your roadmap, start free with $200 in Deepgram credits and run your own audio through Nova-3, Flux, and the Voice Agent API before you commit.

Best use cases for Deepgram

Real-time voice agents

Flux is priced and built specifically for conversational turn-taking (Flux quickstart), and the Voice Agent API handles interruptions and barge-in over a single WebSocket.

How AssemblyAI compares: its Voice Agent API covers the same bundled loop at the same $4.50/hr, but with a single fixed configuration and no standalone TTS to reuse elsewhere.

Multilingual streaming products

Nova-3 multilingual streams at $0.0058/min promotional, and Nova models cover 45+ languages with automatic language detection (models and languages overview).

How AssemblyAI compares: it publishes multilingual streaming at its base $0.15/hr tier, but no Pro-tier multilingual realtime rate is published.

Regulated and on-prem workloads

Self-hosted deployment, HIPAA BAAs, PCI compliance, and an EU data residency endpoint are all published options.

How AssemblyAI compares: cloud API only, per its published materials, with no self-hosted path published.

High-accuracy batch with domain vocabulary

Keyterm prompting boosts recognition of product names and jargon at $0.0013/min, with diarization and smart formatting included in the base batch rate.

How AssemblyAI compares: keyterms prompting is published at +$0.05/hr and diarization at +$0.02/hr on top of a cheaper base rate, and fully configured it still totals less per hour for batch.

Best use cases for AssemblyAI

High-volume async transcription on a budget

At published rates of $0.15/hr to $0.21/hr for batch work, AssemblyAI undercuts Deepgram's $0.258/hr base for pure transcription jobs where real-time delivery does not matter. If you are transcribing podcast archives or recorded meetings at scale, this is a real cost advantage.

Audio analysis pipelines

AssemblyAI's a-la-carte understanding menu (sentiment, topics, content moderation, key phrases, translation, PII audio redaction) is broader than Deepgram's published intelligence lineup, and its LLM-over-audio tooling is a first-class product. Teams whose product is the analysis, not the transcript, may find the richer menu decisive.

Platform overview of Deepgram

Speech-to-text with Nova-3

Nova-3 serves batch and streaming in 45+ languages with diarization, smart formatting, keyterm prompting, and language detection. See the streaming feature overview.

Flux conversational recognition

Flux is streaming recognition built for voice agents, with turn-taking and conversational latency, from $0.0065/min. The getting started guide compares the streaming paths.

Text-to-speech with Aura-2 and Flux TTS

Aura-2 ($0.030/1k characters) and Flux TTS ($0.045/1k characters) are sold standalone. The Flux TTS overview covers the real-time and REST transports.

Voice Agent API

Bundled or bring-your-own tiers from $0.050/min, with interruptions and turn-taking over one WebSocket. Tiers run Standard, Custom, and Advanced, each with BYO-TTS variants.

Audio Intelligence

Summarization, topic detection, sentiment, and intent recognition, with entity detection and redaction billed per minute. See the Audio Intelligence product page.

Deployment and trust

Cloud, a dedicated EU endpoint, and self-hosted deployment, with SOC 2 Type 1 and 2, HIPAA BAAs, GDPR, CCPA, and PCI. The Flux self-hosted guide covers the GPU requirements.

How to choose

If your priority isLean towardBecause
Real-time voice products (agents, live captions, IVR)DeepgramStreaming rates currently below AssemblyAI's Pro realtime tier, plus Flux and a composable agent stack
Cheapest possible batch transcriptionAssemblyAIPublished base rates of $0.15/hr to $0.21/hr beat Deepgram's $0.258/hr, even with add-ons stacked
Text-to-speech or the full voice loopDeepgramStandalone TTS exists, while AssemblyAI's TTS is only inside its agent bundle
Audio analysis depthAssemblyAIBroader published understanding menu
Self-hosting, data residency, regulated industriesDeepgramPublished self-hosted path, EU endpoint, HIPAA BAAs, PCI
Massive concurrent streaming with no capsTest bothAssemblyAI advertises unlimited streams, and Deepgram publishes caps that rise by plan and are negotiable at Enterprise

Switching from AssemblyAI to Deepgram

  1. Map the endpoints. Batch jobs move from AssemblyAI's async transcript flow to a single POST to Deepgram's pre-recorded endpoint, and streaming moves WebSocket to WebSocket. SDKs are available for Python, JavaScript, Go, and .NET.
  2. Translate feature flags. Speaker labels become diarize=true (included on batch). Formatting becomes smart_format=true (included). Custom vocabulary maps to keyterm prompting. PII redaction maps to the redact parameters.
  3. Move webhooks. AssemblyAI's completion webhooks map to Deepgram's callback URLs on async requests. Payload shapes differ, so plan a thin adapter.
  4. Run in parallel. Mirror one to two weeks of production traffic to both APIs and compare word error rate on your own audio, latency at your percentiles, and invoice line items. The $200 credit typically covers this test at no cost.
  5. Cut over by workload. Teams commonly move streaming first, where the price and latency gap is largest, and batch last, or keep batch on the incumbent if the economics favor it. A partial migration is a legitimate end state.

FAQ

Is Deepgram or AssemblyAI cheaper for real-time streaming?

It depends on the model tier. AssemblyAI's Universal-Streaming is published at $0.15/hr ($0.0025/min), lower than Deepgram's Nova-3 monolingual streaming at $0.0048/min (promotional; regular $0.0077/min). AssemblyAI's higher-accuracy Universal-3.5 Pro Realtime is published at $0.45/hr ($0.0075/min), above Deepgram's current Nova-3 rate. Compare the tiers you would actually run, and test accuracy on your own audio before deciding on price alone.

Which is more accurate, Deepgram Nova-3 or AssemblyAI Universal?

Both vendors publish self-reported benchmarks favoring their own models, and independent results vary by audio domain, noise, and language. The only reliable answer for your workload is a side-by-side test on your own audio, and both platforms make that free to run.

Does AssemblyAI offer text-to-speech?

Not as a standalone API. Its published TTS capability is bundled inside its $4.50/hr Voice Agent API. Deepgram sells standalone TTS (Aura-2 at $0.030/1k characters, Flux TTS at $0.045/1k characters) alongside STT and agents.

Can I self-host either platform?

Deepgram publishes self-hosted deployment documentation for Enterprise customers. AssemblyAI does not publish a self-hosted or on-premises option, and its products run as a cloud API.

How hard is it to migrate from AssemblyAI to Deepgram?

Both are REST plus WebSocket APIs with official SDKs, so migration is mostly endpoint and parameter mapping, plus a thin webhook adapter. Most teams mirror traffic to both for one to two weeks before cutting over.

Which should I choose for a voice agent?

Both bundle a voice agent at $4.50/hr. Deepgram additionally publishes BYO tiers down to $0.050/min with your own LLM and TTS, and AssemblyAI publishes one bundled rate. If you want control over the LLM or voices, Deepgram's tiering is built for that.

Both platforms offer free tiers, so the fastest way to settle this comparison is a side-by-side test on your own audio. When you are ready to size it for production volume, talk to the Deepgram sales team about Growth and Enterprise plans.

Related Articles