Deepgram Review 2026: The Developer Voice AI Platform That Is Fast and Flexible, and Not the Most Accurate on Independent Tests
Correction, Oct 4, 2026: an earlier version dated the Flux Multilingual introduction video to May 1; it was published on May 18, 2026. We also added the Aug 12, 2026 general availability of Flux TTS on the managed API, and corrected the early-access date: Flux TTS was already in early access in the Voice Agent by Jul 31, not first released that day.
Deepgram is not an app you open, it is an API you build on: speech-to-text (Nova-3 and the conversation-aware Flux), text-to-speech (new Flux TTS, plus Aura-2) and a Voice Agent API that bundles both with an LLM. In October 2026 it offers $200 of free credit with no card, raised a $130 million Series C at a $1.3 billion valuation in January, and says it passed $100 million in ARR. This review reads Deepgram's pricing page, docs and changelog on Oct 3, 2026, then checks them against the independent Artificial Analysis boards, where Flux has the lowest latency of any streaming model but Nova-3 sits in the lower half on word error rate. It is also plain about the default that your data is retained for model training unless you opt out.

| Ownership | Independent VC-backed (YC W16); $130M Series C led by AVP, Jan 13, 2026 (reported) |
| Free tier | $200 free credit, no card, no expiry, then pay as you go |
| Cheapest paid plan | $0.26/hr Nova-3 batch transcription; Growth from $4K/yr prepaid |
| Independent check | Flux fastest streaming (0.02 s); Nova-3 WER 5.2%, 46th of 56 |
| Review evidence | Product Hunt 4.9 (81); Trustpilot has 2 reviews, not a signal |
Deepgram is one of the best choices for low-latency, production voice pipelines, as long as you test accuracy on your own audio and set your data options on purpose.
"If you are building a voice agent that has to feel instant, Deepgram's Flux is the model other vendors are measured against, and the platform is built for developers: transparent per-minute prices, regional endpoints, even self-hosting. If you only need the most accurate transcript of a recorded file, the independent numbers say to compare before you commit."
My honest read: Deepgram's reputation is speed and developer ergonomics, and the independent boards back the speed part strongly. Artificial Analysis shows Flux with the lowest latency to a final transcript of any streaming model it lists (0.02 seconds after the speaker stops), and Nova-3 as the fastest model in its batch table by speed factor. They do not back the accuracy marketing: on AA's word error rate index Nova-3 lands 46th of 56 batch models and Flux 30th of 32 streaming models. The other thing to know is the data default: Deepgram retains requests to improve its models unless you send an opt-out flag. This is a VERIFIED EDITORIAL REVIEW: Deepgram's pricing page, docs, changelog, docs index, GitHub org and status were checked on Oct 3, 2026, alongside independent leaderboards, Product Hunt and Hacker News. I did not run a controlled hands-on test.
One platform, four APIs: speech-to-text, text-to-speech, a Voice Agent API and audio intelligence.
Deepgram's own description is "a foundational Voice AI platform for developers, offering unified APIs for Speech-to-Text (STT), Text-to-Speech (TTS), and autonomous Voice Agent orchestration", "available as a managed API (US, EU, and Australia endpoints), in-VPC, or self-hosted". The docs now also list an India endpoint (generally available Sep 15, 2026). On the STT side there are two families: Nova-3, the flagship for recorded and live transcription in 50+ languages, and Flux, a conversational model with built-in end-of-turn detection and interruption handling, in English and a 10-language multilingual version. On the TTS side, Flux TTS is new, with Aura-2 and Aura-1 still available. The Voice Agent API joins listening, an LLM of your choice and speaking over one WebSocket. Deepgram is explicit about its boundary: it does not provide "the reasoning LLM".
Ownership: Deepgram is an independent, venture-funded company (about page: founded in 2015 by Scott Stephenson and a teammate from a dark-matter detector project; Y Combinator's W16 batch per Hacker News launch posts). In January 2026 it raised a $130M Series C at a $1.3B valuation and acquired OfOne, a drive-thru voice AI startup, per its newsroom. I found no acquisition of Deepgram itself.
A well-funded voice infrastructure vendor whose growth numbers are its own, with independent speed rankings that favour it.
Jan 13, 2026, led by AVP per SiliconANGLE; also listed on Deepgram's newsroom. A valuation, not revenue.
From the description of Deepgram's Aug 28, 2026 interview video, with "200,000+ developers" and "2,000+ products". Self-reported, undefined.
Quoted in SiliconANGLE's Series C story. A count of clients, not seats or spend.
github.com/deepgram, Oct 3, 2026. Stars measure developer interest, not usage; the JS SDK has 275.
| Date | Event | Source |
|---|---|---|
| 2015 / 2016 | Founded 2015; seed round of $1.8M reported Sep 2016; YC W16 | Deepgram About; TechCrunch via Hacker News |
| 2022-11-29 | $72M Series B announced | Hacker News item 33786715 |
| 2025-06-16 | Voice Agent API generally available | Hacker News item 44291116 |
| 2026-01-13 | $130M Series C at $1.3B; OfOne acquisition | SiliconANGLE; Deepgram newsroom |
| 2026-04-16 / 05-18 | Flux Multilingual (self-hosted release Apr 16; introduction video May 18) | Deepgram changelog; YouTube |
| 2026-08-12 | Flux TTS generally available on the managed API and in the Voice Agent (already in early access there by Jul 31); GA on self-hosted Aug 26 | Deepgram changelog |
| 2026-08-28 | Interview: $100M ARR and Flux TTS launch | Deepgram YouTube |
| 2026-09-15 | India endpoint generally available | Deepgram changelog |
| 2026-10-29 | Deepgram Speak '26 conference, San Francisco | deepgram.com docs index |
A 15-minute theCUBE interview posted on Deepgram's own channel: the source of the ARR and developer-count claims above, and the company's own account of Flux TTS.
Published by the official "Deepgram" YouTube channel (https://www.youtube.com/@Deepgram). A CEO interview, so every number in it is a company claim.
Developers praise speed, diarization and turn detection. Independent review volume is thin.
On a May 5, 2026 thread about OpenAI's low-latency voice stack, a commenter wrote: "People are migrating to the 'End Of Thought' triggers. Deepgram does that wonderfully." On Aug 28, 2026, in a thread about Google's Gemini 3.5 Transcribe, a user said it still lacks real-time diarization beyond three people "when others do it very well, like Soniox and Deepgram". A Feb 2026 post titled "Benchmarking STT providers on real calls (Deepgram 15.9% vs. OpenAI 39.8% WER)" links to a single tweet, so it is one person's test. On a Jul 2026 thread, one commenter noted that a "local-first" app actually used "Deepgram cloud" for transcription.
Reviewers praise "fast, accurate transcription", low latency for real-time products, easy integration and speaker diarization. One builder wrote that Deepgram offered "higher accuracy for diverse accents and technical terminology (95%+ vs 85-90%)" against Google and AWS, which is that builder's own test. Criticisms: debugging transparency when quality drops, text-to-speech that "could be enhanced", and limited visibility into why transcripts are uncertain or delayed. Product Hunt skews toward builders who chose to review.
Trustpilot shows 3.0 from two reviews, both from Aug 2025 and about a different company's sign-up (RAIZR) that the reviewer believes used leaked information. The profile is unclaimed, with no reviews in the past 12 months. This is not evidence about Deepgram's API quality and I do not count it as sentiment, but I report it so the absence of consumer-review volume is visible: Deepgram sells to developers and enterprises, not to people who write reviews.
Pay as you go with $200 free credit, a prepaid Growth plan, and per-minute and per-character rates that you can compute yourself.
Pay As You Go
No minimums, no expiry, no card required.
- All endpoints in public models
- STT concurrency up to 50 REST and 150 WSS
- TTS up to 45, Voice Agent up to 45
- Community and Discord support
Growth
Credits redeemed against usage; up to about 20% off.
- STT up to 50 REST and 225 WSS
- TTS and Voice Agent up to 60
- Overage at your rate plus 10%, billed weekly
- Priority support per Deepgram's FAQ
Enterprise
Contact sales.
- Volume discounts, custom models
- Dedicated or self-hosted (VPC, on-premise)
- BAA for HIPAA, SLAs
Voice Agent API
Standard tier with managed LLM and TTS is $0.075/min ($4.50/hr).
- Advanced tier $0.163/min
- Growth rates about 9 to 22% lower
- Billed per conversation minute
| Product | Pay As You Go | Growth |
|---|---|---|
| Flux English (streaming STT) | $0.0065/min (regular $0.0077) | $0.0057/min (regular $0.0065) |
| Flux Multilingual (streaming STT) | $0.0078/min | $0.0068/min |
| Nova-3 monolingual (streaming) | $0.0048/min (regular $0.0077) | $0.0042/min (regular $0.0065) |
| Nova-3 multilingual (streaming) | $0.0058/min (regular $0.0092) | $0.0050/min (regular $0.0078) |
| Nova-3 monolingual (pre-recorded) | $0.26/hour | $0.22/hour |
| Nova-3 multilingual (pre-recorded) | $0.31/hour | $0.26/hour |
| Whisper Large (pre-recorded, Deepgram-hosted) | $0.29/hour | $0.29/hour |
| Add-ons per minute | Redaction $0.0020; keyterm prompting $0.0013; entity detection $0.0017; diarization $0.0020; smart formatting included | Redaction $0.0017; keyterm $0.0012; entity $0.0017 |
| Flux TTS | $0.0450 per 1,000 characters | $0.0405 |
| Aura-2 TTS | $0.030 per 1,000 characters | $0.027 |
| Aura-1 TTS | $0.0150 per 1,000 characters | $0.0135 |
| Voice Agent: Standard / Standard BYO TTS | $0.075 / $0.065 per min | $0.068 / $0.051 |
| Voice Agent: Custom BYO LLM / BYO LLM + TTS | $0.059 / $0.050 per min | not listed / $0.041 |
| Voice Agent: Advanced / Advanced BYO TTS | $0.163 / $0.122 per min | $0.146 / $0.110 |
| Check you can do | Result (JAVIS arithmetic) |
|---|---|
| Streaming Nova-3 promo versus regular | $0.0048 vs $0.0077 per min: the promo is 37.7% lower; per hour $0.288 vs $0.462 |
| 1,000 hours a month, Nova-3 pre-recorded | 1,000 x $0.26 = $260 (Growth $220) |
| 1,000 hours a month, Nova-3 streaming | 60,000 min x $0.0048 = $288 at the promo rate; x $0.0077 = $462 at the regular rate |
| $200 free credit in Nova-3 streaming minutes | $200 / $0.0048 = 41,667 min (694 hours); Deepgram's FAQ says about 43,000 minutes (over 700 hours), a slightly different figure |
| 10,000 voice-agent minutes a month | Standard $750; BYO LLM and TTS $500; Advanced $1,630 |
| Growth discount, Nova-3 streaming promo | $0.0042 vs $0.0048 = 12.5%; Growth TTS is 10% below Pay As You Go |
Two things to read twice. First, the "limited-time promotional rates on streaming" carry no end date on the page, so budget at the regular rate if you need certainty. Second, the home page still says Flux TTS is free "through September 12th" while the pricing page now charges $0.045 per 1,000 characters, so the home page text is out of date.
Decide by whether your product talks in real time, and where your audio is allowed to go.
Start with Flux STT and the Voice Agent API on the $200 credit. AA shows Flux with the lowest streaming latency.
Benchmark Nova-3 against two rivals on your own audio. On AA's batch index Nova-3 is 5.2% WER against 2.2% to 4.1% for several others.
All are add-ons per minute; diarization and redaction are $0.0020 per minute each on Pay As You Go.
Send
mip_opt_out=true on every request, use a regional endpoint, and ask sales about a BAA, VPC or self-hosting.Check Flux TTS and Aura-2 ($30 to $45 per 1M characters) against cheaper rivals first.
Growth from $4K a year, then Enterprise volume pricing and self-hosted containers.
Flux is speech recognition that knows when a person has finished talking, which is the hard part of a voice agent.
Deepgram's docs describe Flux as "conversational speech recognition purpose-built for voice agents, with built-in end-of-turn detection, natural interruption handling, and word-level timestamps." It runs only on the /v2/listen endpoint, comes in English (flux-general-en) and a 10-language version with in-conversation language switching (flux-general-multi), and for voice agents it replaces the usual pattern of a separate voice-activity detector. Nova-3 remains the choice for recorded audio and for live transcription that does not manage turns, with 50+ languages, keyterm prompting, diarization and smart formatting.

| Independent check (Artificial Analysis, read Oct 3, 2026) | Result | Comment |
|---|---|---|
| Streaming latency to final transcript after speech end | Deepgram Flux 0.02 s, the lowest in AA's highlights; next best Cartesia Ink-2 and Inworld STT 1 at 0.07 s | Supports Deepgram's speed claim |
| Streaming word error rate (AA-WER Streaming) | Nova-3 Realtime 6.6% (27th of 32); Flux 7.4% (30th of 32); best listed 2.5% | Flux trades some accuracy for turn-taking and speed |
| Batch word error rate (AA-WER v2) | Nova-3 5.2% (46th of 56 rows); Whisper Large v3 (fal.ai) and Large v2 (OpenAI) 4.1%; ElevenLabs Scribe v2 2.2%; AssemblyAI Universal-3 Pro 3.1% | Does not support "more accurate than Whisper" on these datasets |
| Batch speed factor | Nova-3 585.9, the highest in the table | Supports the speed claim |
| Batch price per 1,000 minutes | Nova-3 $4.30 (rank 32 of 56 by price) | Middle of the pack |
Here is why I raise it. Deepgram's developer docs say it is "36% more accurate, up to 5x faster, and 1.4x cheaper" than OpenAI Whisper and more accurate than Amazon, Google and Microsoft. On AA's index, Nova-3 is faster by a wide margin but less accurate than the Whisper entries it lists at 4.1% (a Replicate-hosted Whisper Large v3 entry scores 10.1%), and several newer rivals are well ahead. Different datasets, normalization and test dates can explain a lot, and Deepgram's numbers come from its own comparisons. My practical advice: treat Deepgram's accuracy claims as a hypothesis, and test with your own noisy, accented, domain-heavy audio before you pay.
Deepgram's own 4-minute walkthrough of the multilingual Flux model: 10 languages, code-switching, turn detection and interruption handling on one connection.
Published by the official "Deepgram" YouTube channel (https://www.youtube.com/@Deepgram). Published May 18, 2026. Vendor tutorial, so it shows the best case; it links the live playground.
- Lowest streaming latency on the independent board, with turn detection built in.
- True per-second billing, public prices, and concurrency limits published by plan.
- Deployment range: managed API, regional endpoints, VPC and self-hosted containers.
- Accuracy on your audio: AA's WER ranks put Nova-3 and Flux in the lower half.
- Promotional streaming rates with no end date.
- Default data retention unless you opt out.
Audio in, transcript and turn events out, an LLM in the middle, then speech back, with your logic around it.
The review step is not decoration. Voice agents fail in the gaps between components, and Deepgram's own docs devote pages to audio playback, echo cancellation and concurrency handling. Plan to listen to real calls, not demos.

One API for listening, thinking and speaking, billed by the conversation minute, with the LLM tier doing most of the cost.
The Voice Agent API (generally available since June 2025 per Hacker News posts) is a single WebSocket that unifies STT, LLM orchestration and TTS. Pricing is per minute by tier: Standard $0.075, Custom with your own LLM $0.059, Custom with your own LLM and TTS $0.050, Advanced $0.163. The docs list managed LLMs from OpenAI, Anthropic, Google and NVIDIA, with Groq and Amazon Bedrock as bring-your-own endpoints. The changelog added claude-sonnet-5 as an Advanced-tier managed Anthropic model on Jul 1, 2026 and deprecated Claude Sonnet 4, and flagged the Gemini 2.5 Flash family for deprecation in October. Flux TTS, which Deepgram says conditions speech on conversational context with latency "as low as 80 milliseconds", was already in early access in the Voice Agent by Jul 31, 2026 and became generally available on the managed API and in the Voice Agent on Aug 12, 2026; self-hosted general availability followed on Aug 26.

| Job | How Deepgram does it | What is public | What to watch |
|---|---|---|---|
| Real-time phone agent | Voice Agent API with Flux STT, LLM, Flux TTS or Aura-2 | Per-minute tiers, concurrency up to 45 (Growth 60) | Tier choice drives cost: Advanced is 2.2 times Standard |
| Contact-centre transcription | STT plus audio intelligence | Amazon Connect, Genesys, AudioCodes and Twilio guides | Add-ons bill per minute on top of transcription |
| Meeting and podcast transcripts | Nova-3 batch with diarization | $0.26/hour plus $0.0020/min diarization | AA batch WER is mid-table |
| Self-hosted voice stack | Containers in your VPC or on-premise | Self-hosted releases for Flux, Aura-2 and Flux TTS (GA Aug 26, 2026) | Needs NVIDIA GPUs; Enterprise contract |
| Drive-thru and restaurants | Solutions page and the OfOne acquisition | Marketing claims | No public price or case metrics found |
Broad telephony and voice-agent integrations, official SDKs in five languages, and a CLI with a built-in MCP server.
Deepgram lists integrations with Twilio, Vapi, LiveKit, Pipecat, Jambonz, Amazon Connect, Genesys, AudioCodes, Google Dialogflow CX, Zoom and AWS S3, plus IBM watsonx Orchestrate and Cloudflare as partners. For developers, the dg CLI (released Apr 2026) exposes STT, TTS and account management and includes an MCP server, with documented config for Claude Code (dg mcp over stdio). If MCP is new to you, JAVIS's Claude MCP and Connectors guide explains the idea, and the Claude Code setup guide covers the agent that config targets. Competitor Speechify even ships a proxy so Deepgram's Voice Agent can use its voices, a sign of how widely Deepgram is used as the agent layer.
| Integration | What it does | Status as read |
|---|---|---|
| Twilio, Vapi, LiveKit, Pipecat, Jambonz | Telephony and voice-agent frameworks | Named in Deepgram's Flux Multilingual description and docs |
| Amazon Connect, Genesys, AudioCodes, Dialogflow CX | Contact-centre and IVR stacks | Documented guides on developers.deepgram.com |
| Python, JavaScript, Go, .NET, Java SDKs | Official clients | Python 468 stars, JS 275, Go 90, .NET 55, Java 11; MIT; pushed Sep 28 to Oct 2, 2026 |
| dg CLI and MCP server | STT/TTS from the terminal and from MCP clients | MIT, 10 stars, pushed Oct 3, 2026; Claude Code config documented |
| Self-hosted containers | Run in your VPC or on-premise | self-hosted-resources repo, 39 stars; monthly releases |
| Docs MCP server and skills repo | Agent-friendly docs and prompts | Documented at developers.deepgram.com/_mcp/server; skills repo 23 stars |
Deepgram is cheap and clear for speech-to-text, and mid-to-expensive for text-to-speech.
| Monthly use | Deepgram (JAVIS arithmetic, Pay As You Go) | Comparison | Note |
|---|---|---|---|
| 10M characters of TTS | Flux $450; Aura-2 $300; Aura-1 $150 | Murf Falcon 2 $100; Speechify API $91 on Starter (sticker prices, Oct 3, 2026) | Different models and quality; Deepgram's Aura-2 is not on the Artificial Analysis TTS board |
| 1,000 hours batch STT | $260 on Nova-3 mono | AA lists Nova-3 at $4.30 per 1,000 minutes ($258 per 1,000 hours) | Matches within rounding |
| 1,000 hours streaming STT | $288 (promo) to $462 (regular) on Nova-3; $390 to $462 on Flux | No cross-vendor price verified here | Promo has no end date |
| 10,000 agent minutes | $500 to $1,630 depending on tier | Includes STT, LLM and TTS in Deepgram's bill | Bring-your-own LLM lowers the Deepgram bill but you pay the LLM vendor |
| Concurrency | STT 150 WSS; TTS 45; agent 45 on Pay As You Go | Limits apply per project, not per account | Secondary self-serve projects are limited to one concurrent stream |
Strong compliance and deployment options, and a default that retains your data for model improvement.
| Control | What Deepgram says |
|---|---|
| Compliance | SOC 2 Type 1 and Type 2, HIPAA with BAAs for Enterprise, GDPR ready, CCPA, PCI (vendor statements; reports not read) |
| Encryption | TLS 1.3 in transit, AES-256 at rest (Model Improvement page) |
| Data retention default | Requests participate in the Model Improvement Program by default and are retained, including audio, transcripts, TTS text and agent audio |
| Opt-out | mip_opt_out=true per request gives zero data retention beyond processing; check Console Usage > Logs; account-wide enforcement on request |
| Data residency | Regional endpoints in US, EU, Australia and India; fully in-region only when the request is also opted out of model improvement |
| Self-hosting | Containers in your VPC or on-premise (NVIDIA GPUs), FIPS 140-3 images GA Jul 28, 2026; Enterprise |
| Logs | Request metadata and usage logs retrievable for 90 days; they do not contain audio or transcripts |
| Third-party agents | Voice Agent traffic to third-party LLMs follows those providers' terms when you bring your own keys |
mip_opt_out=true, and that data used for training is "only the data included through" that program. It also says it never sells or redistributes it. If your audio contains customer, patient or employee speech, set the flag in code, confirm it in the console log, and put it in your contract rather than relying on remembering it.Best for teams building real-time voice products who value latency and deployment control.
| Situation | Why Deepgram fits | Likely plan |
|---|---|---|
| Startup building a voice agent | Flux turn detection, Voice Agent API, $200 free credit | Pay As You Go |
| Contact-centre or telephony platform | Amazon Connect, Genesys, AudioCodes guides and concurrency | Growth or Enterprise |
| Healthcare or regulated team | BAA, regional endpoints, VPC and self-hosted options | Enterprise |
| Developer using an AI coding agent | dg CLI with MCP server and documented Claude Code config | Pay As You Go |
| Team needing the lowest transcription error rate on recordings | AA ranks Nova-3 46th of 56 on batch WER | Compare ElevenLabs, AssemblyAI and others first |
| Team wanting the cheapest expressive TTS | Aura-2 at $30 per 1M characters is above Murf and Speechify list prices | Compare TTS vendors first |
| Privacy-first team that cannot retain audio | Zero retention needs the opt-out flag on every request | Enterprise or self-hosted |
Six real gaps, not manufactured ones.
On AA, Nova-3 is 5.2% WER (46th of 56 batch rows) and Flux 7.4% (30th of 32 streaming), while Deepgram's own copy claims better accuracy than Whisper.
Model Improvement Program participation is the default; zero retention needs mip_opt_out=true on every request, and regional residency needs both settings.
Streaming rates are "limited-time", the home page still advertises a Flux TTS free period that ended Sep 12, and the FAQ's free-credit minutes differ slightly from the table.
Flux TTS $45 and Aura-2 $30 per 1M characters against $10 for Murf's Falcon 2 and $6.6 for Speechify's Simba on the AA list; Aura-2 is not on the AA board.
Pay As You Go allows 150 WSS STT and 45 TTS or agent streams, and secondary self-serve projects get one stream; higher limits need Growth.
Trustpilot has two unrelated reviews, G2 and Capterra ratings were not available to us, and models and LLMs deprecate quickly (Whisper support removed from self-hosted in Oct 2026, Claude Sonnet 4 deprecated).
Deepgram is best for real-time developer voice pipelines. Others win on accuracy, TTS price or finished apps.
| If you mostly need... | Compare Deepgram with... |
|---|---|
| Top-ranked text-to-speech voices | ElevenLabs review (Eleven v4 is rank 1 on the AA TTS list; ElevenLabs Scribe v2 is third in AA's batch WER table at 2.2%) |
| A finished transcription and notes app | Otter.ai review or Fireflies.ai review |
| Editing audio and video by editing the transcript | Descript review |
| Cheaper TTS per character | Murf's Falcon 2 ($10 per 1M) and Speechify's Simba 3.2 ($6.6 on the AA list), both verified Oct 3, 2026 |
| Wiring voice into coding agents | Claude MCP and Connectors guide and Claude Code setup guide |
| Entry point (verified Oct 3, 2026) | Deepgram STT | Deepgram TTS | Deepgram Voice Agent |
|---|---|---|---|
| Free | $200 credit, no card, no expiry | Same credit | Same credit |
| Cheapest paid | $0.26/hour Nova-3 batch; $0.0048/min streaming promo | $0.015 per 1,000 characters (Aura-1) | $0.050/min BYO LLM and TTS |
| Unit that runs out | Prepaid credit, per second | Per 1,000 characters | Per conversation minute |
| Self-host | Enterprise containers | Flux TTS GA on self-hosted Aug 26, 2026 | Not stated |
No affiliate relationship shapes this review.
JAVIS has not established a publisher affiliate program with Deepgram. Every CTA in this article points to Deepgram's official site. If that changes, the review's facts, limitations and alternatives will not.
Go to the official page.
Open DeepgramRead the alternative that matches your real question.
Choose the next articleDeepgram sits between speech APIs, voice agents and finished transcription apps.
Every card below links to a published JAVIS review.
Card images come from each product's own site or its JAVIS review page.
Questions people are actually searching right now
Is Deepgram free?
Deepgram gives $200 of free credit with no card and no expiry, then charges pay as you go. There is no permanent free tier beyond the credit. Verified on deepgram.com/pricing, Oct 3, 2026.
How much does Deepgram cost in October 2026?
Nova-3 batch transcription is $0.26 an hour, streaming Nova-3 is $0.0048 a minute at the current promotional rate (regular $0.0077), Flux English streaming is $0.0065 a minute, TTS is $0.015 to $0.045 per 1,000 characters, and the Voice Agent API is $0.050 to $0.163 a minute by tier. Growth starts at $4K a year prepaid.
Is Deepgram more accurate than Whisper?
Deepgram says so. On Artificial Analysis's batch index, Nova-3 scored 5.2% word error rate against 4.1% for Whisper Large v3 (fal.ai) and Large v2 (OpenAI), though a Replicate-hosted Large v3 entry scored 10.1%, so the independent board does not support the claim for those datasets. Test on your own audio.
What is Deepgram Flux?
A conversational speech recognition model with built-in end-of-turn detection and interruption handling for voice agents, in English and a 10-language multilingual version. On Artificial Analysis it has the lowest streaming latency listed (0.02 seconds).
Does Deepgram store or train on my audio?
By default yes: requests join the Model Improvement Program and are retained to improve models. Send mip_opt_out=true to opt out and get zero retention, and confirm it in the console logs.
Can I self-host Deepgram?
Yes, for Enterprise customers: containers in your own VPC or on-premise on NVIDIA GPUs, with monthly releases and FIPS 140-3 images.
Does Deepgram work with Claude Code?
The dg CLI includes an MCP server with documented config for Claude Code, and the Voice Agent API offers Anthropic's claude-sonnet-5 as a managed LLM on its Advanced tier.
What are the best Deepgram alternatives?
ElevenLabs and AssemblyAI for transcription accuracy, Murf or Speechify for cheaper TTS, and Otter.ai or Fireflies.ai if you want a finished meeting app. See Alternatives above.
Where this came from, and when it was checked.
Pricing and plans: deepgram.com/pricing, checked Oct 3, 2026 (streaming and pre-recorded tabs, minute and hour toggles, FAQ answers expanded); deepgram.com docs index.
Product and launches: developers.deepgram.com changelog (2026 entries including Jul 1, Jul 31, Aug 12, Aug 26, Sep 15, Oct 1), docs index, rate-limits, Your Data, Model Improvement Partnership Program and CLI MCP pages; deepgram.com About and Newsroom; GitHub data for the deepgram org.
Market: SiliconANGLE (Jan 13, 2026) for the Series C and client count; Hacker News titles for earlier rounds; the Deepgram YouTube video description for the ARR claim (company-stated).
Independent check: artificialanalysis.ai Speech to Text non-streaming and streaming boards, checked Oct 3, 2026, and the Text to Speech board (Aura-2 not listed).
Real users: Product Hunt (4.9, 81 reviews) and Trustpilot (3.0, 2 reviews); Hacker News. G2, Capterra and Reddit ratings were not available to us, so none is cited.
Official videos: youtube.com/watch?v=0ZGsq6HSIvw and QgQQL03bbTk; author "Deepgram" on YouTube.
JAVIS Fit Score (77/100): editorial composite of latency and developer ergonomics (strong), pricing clarity and deployment range (strong), independent accuracy rank (weak), default data retention and promotional pricing (mixed), TTS price (mixed), thin independent reviews (weak). Weighted by judgement; no automated test was run.




