Back to Blog
Provider Integrations

Deepgram Flux vs AssemblyAI Universal-3.6 Pro for Voice Agents

Compare current Deepgram Flux and AssemblyAI Universal-3.6 Pro streaming models, scoped benchmark evidence, languages, costs and their Burki setup.

Meeran Malik
8 min read

A useful Deepgram vs AssemblyAI comparison starts with the work the model will do. For a live voice agent, recognizing the caller's words, deciding when a turn ends and preserving a corrected phone number matter together. A transcription score from an unrelated recording workload does not settle that decision.

As checked October 6, 2026, the relevant models here are Deepgram Flux, including flux-general-en and flux-general-multi, and AssemblyAI Universal-3.6 Pro Realtime, selected as universal-3-6-pro. AssemblyAI announced 3.6 on September 29. Its streaming documentation now names 3.6 as the default when no model is supplied, while 3.5 remains available. The provider's async flagship has a different version, so “latest AssemblyAI” needs a task qualifier. AssemblyAI changelog, streaming model selection.

The comparison below separates vendor benchmarks from configuration choices. Burki did not run paid API tests or real calls to create this article. Download the free evaluation worksheet to plan a fair check of your own workflow.

Compare the exact streaming products

Flux has its own /v2/listen endpoint and conversation-oriented turn events. English and multilingual are different model IDs; the multilingual family covers ten languages. It is not a Nova-2 alias, and /v1/listen is not the Flux endpoint. Deepgram Flux quickstart.

DecisionDeepgram FluxAssemblyAI Universal-3.6 Pro
Exact streaming modelflux-general-en or flux-general-multiuniversal-3-6-pro
Language scopeEnglish model; ten-language multilingual model32-language streaming model
Conversation behavior to evaluateIntegrated turn events, eager-turn behavior, interruption recoveryTurn completion, contextual recognition and corrected entities
Current Burki selectionDeepgram recognition with the appropriate Flux modelAssemblyAI (BYO), with your own provider account
Separate model choice in GPT Live?NoNo

Check the current AssemblyAI language list and selection rules before deciding whether code switching fits your callers. The Flux language list includes English, Spanish, German, French, Hindi, Russian, Portuguese, Japanese, Italian and Dutch. Arabic is not on that list. A Dubai business cannot infer Arabic support from the word “multilingual.”

The original Flux English launch was October 2, 2025; multilingual became generally available April 29, 2026. Those are existing products being compared with a recent AssemblyAI version, not two models launched today. Flux English release, multilingual announcement.

What the available benchmark actually says

AssemblyAI publishes a vendor-run voice-agent corpus of 12,460 scripted recordings across three noise environments, three utterance lengths and twelve speaker groups. Its default streaming harness fixes language and applies the same normalizer. In that report, normalized entity error rate, or NEER, measures incorrectly transcribed entities. Lower is better. It is not ordinary word error rate. AssemblyAI benchmark report.

Vendor-reported resultUniversal-3.6 ProFlux EN
Overall NEER14.4%30.1%
Phone-entity error2.4%11.5%

The table's run span is August 27 to September 22, 2026. Its per-model registry gives September 3–21 for 3.6 and August 19–September 11 for Flux EN. These are different run dates, including prelaunch 3.6 measurements, on a vendor-controlled corpus. Competitor defaults may improve with tuning. The report also has an older “data as of” footer; use its table-specific windows rather than pretending the whole page represents a single October test. Benchmark methodology and run registry.

Those results support a closer look at entity recognition for this test. They do not establish the accuracy of your street names, language mix, phone codec or configured Burki agent.

The independent Pipecat/Daily benchmark checked October 6 includes 1,000 samples from smart-turn-data-v3.1-train, with semantic WER and time to final segment. Its current summary has Nova-3-general and Universal-3.6 Pro rows, but no Flux row. It therefore cannot supply an independent Flux-versus-3.6 winner. Its September scoring update also changed error counting, so older copied scores need the original scoring version. Original Pipecat benchmark.

Keep accuracy and latency separate

A caller saying “no” incorrectly transcribed as “go” is an accuracy failure. A correct number cut off halfway through is a turn-boundary failure. A correct transcript followed by a slow external lookup is a different delay again. Treating all three as “the STT is bad” makes model choice harder.

Deepgram's quickstart describes roughly 260ms end-of-turn detection and configurable thresholds. That is a vendor description of a component boundary, not an independently comparable end-to-end call measurement. Its eager-turn mode can prepare a reply early, but must handle the caller resuming speech. Flux turn configuration.

Measure the same boundary for both candidates: caller end of speech to final transcript. Separately record the first audible agent response. Do not compare one provider's time to first partial with another's final transcript, or add the fastest component measurements from different calls.

Illustrative comparison checklist showing identical caller audio, fixed model settings, transcript and entity scoring, then separate audible response review.

Illustrative evaluation flow, not a benchmark performed by Burki. Keep the input fixed and judge the caller's actual task.

Price the same workload and billing unit

The provider pages checked October 6 list these USD pay-as-you-go rates:

ProductDisplayed provider rateArithmetic for 60 billed minutes
Flux EnglishPromotional USD0.0065/min; regular USD0.0077/minUSD0.39 promotional; USD0.462 regular
Flux MultilingualUSD0.0078/minUSD0.468
Universal-3.6 Pro RealtimeUSD0.45/hourUSD0.45

The calculations simply multiply the displayed rates; they are not a Burki quote or a real invoice. The Flux English promotion's end date was not established in this review, so confirm the rate in your account before budgeting. Deepgram pricing, AssemblyAI pricing.

Billing quantity can change the comparison. AssemblyAI says streaming is charged for the whole open WebSocket session, including time without submitted audio, until the session ends. Deepgram's pricing page describes processed audio duration. A minute of listening audio and a minute of open session are not automatically the same workload. AssemblyAI session billing, Deepgram pricing.

Add language-model generation, speech synthesis, Burki charges and telephone costs separately. See Burki pricing for the platform context. A small per-minute difference should not override a model's failure on the critical fields your business needs.

Configure the candidates in Burki

The current deployed interface and audited source include both Flux model IDs and AssemblyAI 3.5/3.6. That confirms a configuration path, not acceptance of your provider credentials or telephone route.

For a separate recognition comparison, open the assistant's Voice settings and choose Standard pipeline · advanced provider choices under Conversation engine. If that control is collapsed, expand Model and provider choices to find it. Then expand Model and provider choices for Advanced speech recognition, and use Recognition provider and Recognition model.

For Deepgram, select the appropriate Flux model and a supported language configuration. Begin with the existing turn settings; change one threshold only after a repeatable failure points to turn timing.

For AssemblyAI, select AssemblyAI (BYO) and universal-3-6-pro. Set speech recognition to BYO in Usage & billing and supply your authorized AssemblyAI API key. Managed AssemblyAI speech is currently unavailable. The optional Speech recognition prompt accepts up to 1,750 characters and describes the recognition domain; it is separate from the assistant's conversation instructions. A short example is “Calls to an equipment supplier. Callers provide product identifiers and request staff callbacks.” The current interface also states that the assistant's latest spoken reply is automatically shared with AssemblyAI as context for recognizing the next caller turn. Review that data flow when deciding whether this provider fits your business.

Do not assume every native provider option is available in Burki. For example, choosing 3.6 does not expose every voice-focus, diarization or endpointing control from AssemblyAI's own API. GPT Live handles recognition itself, so changing these STT selectors is not how you change its conversation recognition.

Choose with one representative task

Use fictional reference data first. Include a refusal, a spelled name, a corrected digit, a long pause inside a number, a target-language request and a consented noisy version. Keep source audio, expected values, route, prompts and model IDs in the worksheet.

An AI receptionist collecting callbacks needs the final confirmed number and correct intent. A low average WER is helpful evidence, but it cannot compensate for losing the one digit that makes the callback impossible. Pick the model that passes your defined workflow after authorized testing, record unresolved cases and retain a human fallback.

Ready to try Burki?

Create an assistant and check your available browser practice allowance.

Create your assistant

Trial eligibility and available practice are shown in your workspace.

Related Articles