Deepgram Flux vs AssemblyAI Universal-3.6 Pro for Voice Agents
Compare current Deepgram Flux and AssemblyAI Universal-3.6 Pro streaming models, scoped benchmark evidence, languages, costs and their Burki setup.
Table of Contents▼
A useful Deepgram vs AssemblyAI comparison starts with the work the model will do. For a live voice agent, recognizing the caller's words, deciding when a turn ends and preserving a corrected phone number matter together. A transcription score from an unrelated recording workload does not settle that decision.
As checked October 6, 2026, the relevant models here are Deepgram Flux, including flux-general-en and flux-general-multi, and AssemblyAI Universal-3.6 Pro Realtime, selected as universal-3-6-pro. AssemblyAI announced 3.6 on September 29. Its streaming documentation now names 3.6 as the default when no model is supplied, while 3.5 remains available. The provider's async flagship has a different version, so “latest AssemblyAI” needs a task qualifier. AssemblyAI changelog, streaming model selection.
The comparison below separates vendor benchmarks from configuration choices. Burki did not run paid API tests or real calls to create this article. Download the free evaluation worksheet to plan a fair check of your own workflow.
Compare the exact streaming products
Flux has its own /v2/listen endpoint and conversation-oriented turn events. English and multilingual are different model IDs; the multilingual family covers ten languages. It is not a Nova-2 alias, and /v1/listen is not the Flux endpoint. Deepgram Flux quickstart.
| Decision | Deepgram Flux | AssemblyAI Universal-3.6 Pro |
|---|---|---|
| Exact streaming model | flux-general-en or flux-general-multi | universal-3-6-pro |
| Language scope | English model; ten-language multilingual model | 32-language streaming model |
| Conversation behavior to evaluate | Integrated turn events, eager-turn behavior, interruption recovery | Turn completion, contextual recognition and corrected entities |
| Current Burki selection | Deepgram recognition with the appropriate Flux model | AssemblyAI (BYO), with your own provider account |
| Separate model choice in GPT Live? | No | No |
Check the current AssemblyAI language list and selection rules before deciding whether code switching fits your callers. The Flux language list includes English, Spanish, German, French, Hindi, Russian, Portuguese, Japanese, Italian and Dutch. Arabic is not on that list. A Dubai business cannot infer Arabic support from the word “multilingual.”
The original Flux English launch was October 2, 2025; multilingual became generally available April 29, 2026. Those are existing products being compared with a recent AssemblyAI version, not two models launched today. Flux English release, multilingual announcement.
What the available benchmark actually says
AssemblyAI publishes a vendor-run voice-agent corpus of 12,460 scripted recordings across three noise environments, three utterance lengths and twelve speaker groups. Its default streaming harness fixes language and applies the same normalizer. In that report, normalized entity error rate, or NEER, measures incorrectly transcribed entities. Lower is better. It is not ordinary word error rate. AssemblyAI benchmark report.
| Vendor-reported result | Universal-3.6 Pro | Flux EN |
|---|---|---|
| Overall NEER | 14.4% | 30.1% |
| Phone-entity error | 2.4% | 11.5% |
The table's run span is August 27 to September 22, 2026. Its per-model registry gives September 3–21 for 3.6 and August 19–September 11 for Flux EN. These are different run dates, including prelaunch 3.6 measurements, on a vendor-controlled corpus. Competitor defaults may improve with tuning. The report also has an older “data as of” footer; use its table-specific windows rather than pretending the whole page represents a single October test. Benchmark methodology and run registry.
Those results support a closer look at entity recognition for this test. They do not establish the accuracy of your street names, language mix, phone codec or configured Burki agent.
The independent Pipecat/Daily benchmark checked October 6 includes 1,000 samples from smart-turn-data-v3.1-train, with semantic WER and time to final segment. Its current summary has Nova-3-general and Universal-3.6 Pro rows, but no Flux row. It therefore cannot supply an independent Flux-versus-3.6 winner. Its September scoring update also changed error counting, so older copied scores need the original scoring version. Original Pipecat benchmark.
Keep accuracy and latency separate
A caller saying “no” incorrectly transcribed as “go” is an accuracy failure. A correct number cut off halfway through is a turn-boundary failure. A correct transcript followed by a slow external lookup is a different delay again. Treating all three as “the STT is bad” makes model choice harder.
Deepgram's quickstart describes roughly 260ms end-of-turn detection and configurable thresholds. That is a vendor description of a component boundary, not an independently comparable end-to-end call measurement. Its eager-turn mode can prepare a reply early, but must handle the caller resuming speech. Flux turn configuration.
Measure the same boundary for both candidates: caller end of speech to final transcript. Separately record the first audible agent response. Do not compare one provider's time to first partial with another's final transcript, or add the fastest component measurements from different calls.
Illustrative evaluation flow, not a benchmark performed by Burki. Keep the input fixed and judge the caller's actual task.
Price the same workload and billing unit
The provider pages checked October 6 list these USD pay-as-you-go rates:
| Product | Displayed provider rate | Arithmetic for 60 billed minutes |
|---|---|---|
| Flux English | Promotional USD0.0065/min; regular USD0.0077/min | USD0.39 promotional; USD0.462 regular |
| Flux Multilingual | USD0.0078/min | USD0.468 |
| Universal-3.6 Pro Realtime | USD0.45/hour | USD0.45 |
The calculations simply multiply the displayed rates; they are not a Burki quote or a real invoice. The Flux English promotion's end date was not established in this review, so confirm the rate in your account before budgeting. Deepgram pricing, AssemblyAI pricing.
Billing quantity can change the comparison. AssemblyAI says streaming is charged for the whole open WebSocket session, including time without submitted audio, until the session ends. Deepgram's pricing page describes processed audio duration. A minute of listening audio and a minute of open session are not automatically the same workload. AssemblyAI session billing, Deepgram pricing.
Add language-model generation, speech synthesis, Burki charges and telephone costs separately. See Burki pricing for the platform context. A small per-minute difference should not override a model's failure on the critical fields your business needs.
Configure the candidates in Burki
The current deployed interface and audited source include both Flux model IDs and AssemblyAI 3.5/3.6. That confirms a configuration path, not acceptance of your provider credentials or telephone route.
For a separate recognition comparison, open the assistant's Voice settings and choose Standard pipeline · advanced provider choices under Conversation engine. If that control is collapsed, expand Model and provider choices to find it. Then expand Model and provider choices for Advanced speech recognition, and use Recognition provider and Recognition model.
For Deepgram, select the appropriate Flux model and a supported language configuration. Begin with the existing turn settings; change one threshold only after a repeatable failure points to turn timing.
For AssemblyAI, select AssemblyAI (BYO) and universal-3-6-pro. Set speech recognition to BYO in Usage & billing and supply your authorized AssemblyAI API key. Managed AssemblyAI speech is currently unavailable. The optional Speech recognition prompt accepts up to 1,750 characters and describes the recognition domain; it is separate from the assistant's conversation instructions. A short example is “Calls to an equipment supplier. Callers provide product identifiers and request staff callbacks.” The current interface also states that the assistant's latest spoken reply is automatically shared with AssemblyAI as context for recognizing the next caller turn. Review that data flow when deciding whether this provider fits your business.
Do not assume every native provider option is available in Burki. For example, choosing 3.6 does not expose every voice-focus, diarization or endpointing control from AssemblyAI's own API. GPT Live handles recognition itself, so changing these STT selectors is not how you change its conversation recognition.
Choose with one representative task
Use fictional reference data first. Include a refusal, a spelled name, a corrected digit, a long pause inside a number, a target-language request and a consented noisy version. Keep source audio, expected values, route, prompts and model IDs in the worksheet.
An AI receptionist collecting callbacks needs the final confirmed number and correct intent. A low average WER is helpful evidence, but it cannot compensate for losing the one digit that makes the callback impossible. Pick the model that passes your defined workflow after authorized testing, record unresolved cases and retain a human fallback.
Ready to try Burki?
Create an assistant and check your available browser practice allowance.
Create your assistantTrial eligibility and available practice are shown in your workspace.