Back to Blog
ROI & Business

Scaling voice AI: plan concurrency, limits and recovery

Estimate peak simultaneous sessions, inspect every dependency and expand only after bounded load and outcome checks.

Burki
(Updated: September 25, 2026)
2 min read

Monthly call volume is not the same as peak capacity. Ten thousand calls spread across a month can require less simultaneous capacity than a short burst of a few hundred calls. Plan for arrival patterns, call duration and dependent-system limits together.

Estimate concurrency first

A rough steady-state planning estimate is arrival rate multiplied by average call duration, using matching time units. For example, two arrivals per minute with a three-minute average duration imply about six active calls on average. This is an illustrative average, not a safe peak-capacity guarantee. Bursts, long calls and transfers require additional analysis.

Inspect each limiting resource

ResourceWhat to verify
CarrierConcurrent sessions, call-start limits and route restrictions
Voice/model providerApplicable account quotas and request/audio limits
Worker and media serviceAdmission capacity, startup behavior and cleanup
Business actionsCalendar, CRM or database request limits
FundingMaximum reservation exposure and spend caps
Human fallbackActual receiving capacity and unavailable behavior

The lowest effective limit can constrain the workflow. Do not assume capacity automatically expands because an underlying provider advertises large deployments.

Test a bounded increase

Begin after a sequential scenario works. Increase concurrency in permitted steps, use controlled endpoints and agree on stop conditions and maximum expense. Measure response timing, failed admissions, action failures and post-hangup cleanup. Follow the load-test plan.

Preserve configured caps until an authorized decision changes them. A rejected admission can be the correct behavior when capacity or funds are exhausted; inspect its reason instead of treating every rejection as a defect.

Monitor outcomes at volume

Compare the same task at each tested load. Include failures and unresolved actions in the sample. Automated evaluations can help prioritize review, but customer satisfaction requires customer feedback and actual business outcomes require evidence.

Burki's architecture uses LiveKit for media with separate application admission and billing. This is not a promise of unlimited simultaneous calls, constant latency or automatic provider failover. State the capacity actually tested on the deployed configuration.

Revisit the cost model

Higher volume can increase total exposure even when an account has a lower unit rate. Do not assume a negotiated discount exists. Include provider usage, carrier legs, recurring resources and the staff work needed to investigate failures.

Ready to try Burki?

Create an assistant and check your available browser practice allowance.

Create your assistant

Trial eligibility and available practice are shown in your workspace.

Related Articles