Voice infrastructure
Tune vLLM for voice latency, not just tokens per second
A voice-serving configuration should maximize successful responses within a latency budget. Aggregate tokens per second alone can reward a configuration that makes each caller wait longer. Our short-response sweep increased estimated output from 803.2 to 2,103.0 units per second as concurrency rose from 12 to 48, while p95 first-content time increased from 138.1 to 316.5 ms.
Choose a success metric before a batch size
Suppose your application assigns the LLM a 250 ms p95 first-content budget. This is an example target, not a Mirai service guarantee. In our archived sweep, 32 concurrent requests produced 248.0 ms pooled p95; 48 produced 316.5 ms. The higher setting generated more aggregate output but missed that illustrative budget.
The 32-request observation sits too close to the boundary to establish production headroom. Sustained arrivals, long histories, tool responses and noisy neighbors were not part of that test. Use it to choose candidates for another experiment, not to set a production limit directly.
Define goodput as valid responses meeting the chosen latency and quality conditions per unit of test time. That connects serving capacity to the conversation you want people to have. The archive lacks output-quality scores, so it does not establish goodput.
Read throughput without changing its meaning
The Gemma case study includes 1,520 measured requests across six concurrency levels, with five repetitions per level and two measured waves per repetition. Outputs were capped at 80 tokens and usually much shorter.
At concurrency 48, the saved completion count was 11,373 across 5.408 measured seconds: approximately 2,103.0 recorded output units per second. This excludes pauses between waves. Because the harness could substitute chunk counts for missing token usage, call it an estimated aggregate output rate. It is not 2,103 tokens per second delivered to one caller.
For your next benchmark, retain tokenizer-backed usage for each response and record the entire wall-clock interval, including request arrivals. Keep timeout counts and partial outputs. Pool individual latencies when calculating percentiles; averaging five p95 values does not calculate a pooled p95.
Separate the scheduler controls
| Control | What you are choosing | What to measure |
|---|---|---|
max_num_seqs |
Number of sequences the scheduler can handle together | Queueing, tail latency, request-state pressure |
max_num_batched_tokens |
Token work allowed in a scheduling iteration | Prefill progress and decode intervals |
max_model_len |
Maximum sequence length the deployment accepts | Long-history behavior and rejected requests |
gpu_memory_utilization |
Fraction of device memory budgeted to the engine | Startup allocation and headroom under load |
vLLM’s chunked prefill prioritizes decode work and uses remaining token budget for prefill. Smaller batch-token budgets can improve inter-token latency; larger ones can improve first-token time by advancing prefills faster. Insufficient KV capacity can trigger preemption and recomputation. Those trade-offs are described in the vLLM tuning documentation.
Do not change every control at once. A faster median after increasing memory and changing prompt length does not isolate which change helped.
Use our current recipe as a starting point
These are selected flags from our current configuration, reviewed September 17, 2026, not a complete installation command or the exact archived July launch manifest:
checkpoint: nvidia/Gemma-4-26B-A4B-NVFP4
runtime pin: vLLM 0.26.0
--kv-cache-dtype fp8
--enable-prefix-caching
--enable-auto-tool-choice
--tool-call-parser gemma4
--kernel-config '{"moe_backend":"cutlass"}'
Install a compatible driver and runtime, pin the checkpoint revision and verify the selected kernels on your hardware. Check the official Gemma recipe for version-dependent support. Test ordinary speech, a structured tool call and the batch shapes you expect before declaring a new worker ready. A health endpoint alone does not prove those paths are warm or correct.
Run a workload matrix that can reject a bad change
Include short and long histories, reusable and unique prefixes, fast acknowledgements and longer tool results. Test both fresh and reused connections. Then drive a sustained arrival schedule so the queue can grow; synchronized waves cannot tell you how a service behaves when demand keeps arriving.
For each candidate, save first useful output, token intervals, queue time, request length, failures, cancellations, memory pressure and task results. On interrupted calls, cancel obsolete generation promptly and verify it does not keep consuming capacity. Report the quality and latency thresholds alongside any capacity number.
The community kit contains the observed data and an experiment worksheet. It recomputes the published results without pretending that analysis replay is a fresh GPU benchmark.
Common questions
Does max_num_seqs equal the number of phone calls a GPU supports?
No. It limits sequences scheduled together. Calls generate intermittently, may use tools and audio models, and can have very different context lengths. Capacity requires a realistic arrival-rate and full-stack test.
Should I reduce max_num_batched_tokens to improve every latency metric?
No. With chunked prefill, a smaller budget can protect decode intervals while making prompt processing take more scheduling steps. Tune first-output time and inter-token delay together.
Are the published output rates sustained production throughput?
No. They are estimates from short measured burst waves, excluding pauses and warmup. The legacy counter could fall back from server token usage to content-chunk counts.
Sources and API references
References reviewed on . Measurement dates and experimental limits are stated in the article. Check current API and runtime documentation when integrating.