Scaling voice AI: plan concurrency, limits and recovery
Estimate peak simultaneous sessions, inspect every dependency and expand only after bounded load and outcome checks.
Table of Contents▼
Monthly call volume is not the same as peak capacity. Ten thousand calls spread across a month can require less simultaneous capacity than a short burst of a few hundred calls. Plan for arrival patterns, call duration and dependent-system limits together.
Estimate concurrency first
A rough steady-state planning estimate is arrival rate multiplied by average call duration, using matching time units. For example, two arrivals per minute with a three-minute average duration imply about six active calls on average. This is an illustrative average, not a safe peak-capacity guarantee. Bursts, long calls and transfers require additional analysis.
Inspect each limiting resource
| Resource | What to verify |
|---|---|
| Carrier | Concurrent sessions, call-start limits and route restrictions |
| Voice/model provider | Applicable account quotas and request/audio limits |
| Worker and media service | Admission capacity, startup behavior and cleanup |
| Business actions | Calendar, CRM or database request limits |
| Funding | Maximum reservation exposure and spend caps |
| Human fallback | Actual receiving capacity and unavailable behavior |
The lowest effective limit can constrain the workflow. Do not assume capacity automatically expands because an underlying provider advertises large deployments.
Test a bounded increase
Begin after a sequential scenario works. Increase concurrency in permitted steps, use controlled endpoints and agree on stop conditions and maximum expense. Measure response timing, failed admissions, action failures and post-hangup cleanup. Follow the load-test plan.
Preserve configured caps until an authorized decision changes them. A rejected admission can be the correct behavior when capacity or funds are exhausted; inspect its reason instead of treating every rejection as a defect.
Monitor outcomes at volume
Compare the same task at each tested load. Include failures and unresolved actions in the sample. Automated evaluations can help prioritize review, but customer satisfaction requires customer feedback and actual business outcomes require evidence.
Burki's architecture uses LiveKit for media with separate application admission and billing. This is not a promise of unlimited simultaneous calls, constant latency or automatic provider failover. State the capacity actually tested on the deployed configuration.
Revisit the cost model
Higher volume can increase total exposure even when an account has a lower unit rate. Do not assume a negotiated discount exists. Include provider usage, carrier legs, recurring resources and the staff work needed to investigate failures.
Ready to try Burki?
Create an assistant and check your available browser practice allowance.
Create your assistantTrial eligibility and available practice are shown in your workspace.