Voice infrastructure
Making Gemma 4 26B A4B fast enough for a voice conversation
For a voice agent, optimize the time to useful output at your expected load. Our archived Gemma serving sweep recorded 138.1 ms p95 time to first content at 12 concurrent requests and 316.5 ms at 48. Those are client-side LLM measurements from short, warmed burst tests—not the caller’s full wait for speech. Here is the evidence and what it changes about deployment.
Start with the latency boundary
The useful target is the earliest correct response the caller can hear. A model can generate quickly while the application waits on a new connection, a tool, a sentence boundary or an audio buffer. Measure those separately before changing the GPU.
Our July harness started its timer before opening a TCP connection and stopped the first-output timer at the first nonempty streamed content delta. Role markers and empty deltas did not count. Tool calls and reasoning were not measured as useful output by this harness. Measure the full audio path separately, from the end of the caller’s spoken turn to the first audible reply.
What we ran and what the archive establishes
Our deployment configuration uses nvidia/Gemma-4-26B-A4B-NVFP4, served as voice-agent. The model card lists 25.2 billion total parameters and 3.8 billion active parameters. The active count describes computation, not the full resident weights. This is the 26B A4B variant, distinct from Gemma E4B. See the checkpoint card.
The archived July 28, 2026 sweep used an RTX PRO 6000 Blackwell with 96 GB memory. Each concurrency level has five repetitions. Each repetition sent three waves, discarded the warmup wave, and measured two. The shared system prompt contained 3,403 characters; recorded prompt samples were 758 tokens. Generation used temperature 0.4 and an 80-token output cap. Actual recorded output averaged about 24 units per request.
Deployment notes associate the saved alias with the NVFP4 checkpoint. The individual result files do not archive an immutable checkpoint revision, driver version or complete launch configuration. Our current bootstrap pins vLLM 0.26.0, but that is not proof of the exact July runtime. Treat this as an archived engineering case study with reproducible analysis, not an independently replicated model leaderboard.
The concurrency trade-off in numbers
Percentiles below pool the saved individual request timings across all five repetitions. Throughput divides total recorded completion counts by total measured wave time; warmup and pauses are excluded.
| Concurrent requests | Measured requests | First content p50 | First content p95 | Recorded output/s, estimate |
|---|---|---|---|---|
| 12 | 120 | 127.2 ms | 138.1 ms | 803.2 |
| 16 | 160 | 154.4 ms | 174.2 ms | 980.1 |
| 20 | 200 | 177.8 ms | 195.1 ms | 1,007.2 |
| 24 | 240 | 198.5 ms | 243.9 ms | 1,291.7 |
| 32 | 320 | 213.8 ms | 248.0 ms | 1,515.6 |
| 48 | 480 | 271.0 ms | 316.5 ms | 2,103.0 |
All 1,520 measured requests returned content without a harness-reported error. Output correctness was not scored. Increasing concurrency from 12 to 48 raised estimated output throughput about 2.62 times and p95 first-content latency about 2.29 times. We did not measure the saturation point beyond 48.

Throughput limitation: the legacy harness preferred server-reported completion tokens but fell back to counting content chunks when usage was absent. It did not save which source each request used. Its SSE parser could also skip a line split across reads. These output rates are recorded estimates; they are not audited tokenizer counts or per-user decode speed. The measurement notes preserve these limitations.
Connection reuse belongs in the performance plan
In a separate September 8 diagnostic, 12 serial requests spaced seven seconds apart reused one connection. First useful output ranged from 81.8 to 155.9 ms; the median was 95.35 ms. The server keep-alive timeout was 120 seconds. This probe used a different timing implementation and had no matched concurrent-load baseline, so do not compare it directly with the July table.
The practical change is testable: reuse the inference client through a call, check idle timeouts across every proxy, and record connection creation separately from generation. Preparing a connection while the greeting plays can move setup work off the next turn’s critical path. It does not eliminate inference time.
What we would tune first
Keep a stable prompt prefix, bound context and output, warm the paths you actually use, and choose concurrency against a tail-latency budget. Test thinking behavior explicitly through the model’s chat template; removing a reasoning parser is not the same as disabling reasoning generation. Consult the Gemma serving recipe.
Then run matched precision experiments. This archive has no BF16-versus-NVFP4 latency ablation, so it cannot assign a speedup to quantization alone. The FP4 guide explains the mechanism; the vLLM guide turns these observations into a tuning procedure.
Download and reanalyze the evidence
Download the community kit ZIP, or get the observations JSON, summary CSV and analysis script separately. Put the JSON and script in the same directory, then run python3 analyze.py. No API key, GPU or Python dependency is required to recompute the tables.
Common questions
Is Gemma 4 26B A4B the fastest LLM for voice agents?
These observations do not establish a fastest model. They describe one serving setup and workload. Compare models under the same prompts, quality criteria, hardware, network and arrival rate.
Does 138.1 ms include speech recognition and speech synthesis?
No. It is pooled p95 time from a client request starting to its first nonempty content delta, including a fresh TCP connection. It excludes speech recognition, synthesis and playback.
Can I reproduce the published calculations?
Yes. The downloadable kit includes sanitized per-request first-content timings, aggregate counts and a standard-library Python analysis script. Exact inference replay is limited by missing archived checkpoint revisions and launch manifests.
Sources and API references
References reviewed on . Measurement dates and experimental limits are stated in the article. Check current API and runtime documentation when integrating.