Voice infrastructure
Manage voice-agent context before buying more inference capacity
Keep stable instructions at the beginning, store authoritative facts outside the transcript, and send only the history and tool data needed for the current turn. In our archived RTX 5090 tests, shared-prefix prompts recorded 225.7 ms p95 first-content time versus 501.2 ms for prompts with a unique leading nonce. That is an observed comparison, not a universal cache speedup.
What the shared-prefix test measured
On July 27, 2026, we ran three repetitions per condition on an RTX 5090 with 32 GB memory. Each repetition measured two waves of 16 concurrent requests after a warmup wave. Each condition therefore contains 96 measured requests, with no harness-reported errors.
| Condition | Recorded prompt-token sample | First content p50 | First content p95 |
|---|---|---|---|
| Shared system prefix | 758 | 200.2 ms | 225.7 ms |
| Unique nonce before system prefix | 782 | 359.7 ms | 501.2 ms |
The pooled p95 difference is 275.5 ms, about 55.0% lower in the shared-prefix condition. The unique condition also added 24 prompt tokens. The conditions ran sequentially, and server cache-hit telemetry was not archived. This is useful evidence for testing prompt layout, not an isolated estimate of caching’s causal effect. Both conditions used the legacy measurement harness described in the Gemma case study.
vLLM reuses the KV state of matching prefixes to avoid repeated prefill computation. That benefit concerns prompt processing, not decoding new output tokens. See automatic prefix caching.
Put stable material before changing material
Prefer a stable arrangement of role instructions, voice style and tool definitions. Put changing call identifiers, timestamps and current task state later when your message format permits it. Avoid an unnecessary unique nonce at the very beginning of every prompt.
Stable prefix
Role, voice behavior, stop conditions, common tool schemas
Task context
Current step, allowed actions, validated facts and revisions
Recent conversation
Relevant user turns, assistant replies and paired tool results
Current user turn
The model still receives the same required instructions. This is a layout choice, not permission to omit user-specific facts or cross tenant boundaries. Enforce data access and cache isolation in the serving system. Verify that the rendered token prefix actually stays stable; visually similar strings can tokenize differently.
Context capacity is a budget, not a marketing window
A July 28 startup log reported 67.8 GiB available for KV cache, 645,178 equivalent cache tokens and maximum concurrency of 78.76 at 8,192 tokens per request. These are engine-reported capacity estimates, not a load test at full context.
The arithmetic 645,178 / 8,192 = 78.76 explains the log. Dividing the same budget by 32,768 gives about 19.69. The latter is an illustrative budget calculation only: hybrid attention, allocation groups, shared blocks and runtime overhead mean you must re-profile the actual configuration. Neither division proves an achievable live-call count.
Similarly, setting max_num_seqs to 96 does not promise that 96 full-length histories fit. The historical log used an 8,192-token limit; our current bootstrap default is 32,768. Do not copy the old capacity estimate onto that newer configuration.
An explicit context budget you can adapt
Here is an illustrative 4,096-token application budget, not a measured optimum:
| Component | Token allowance |
|---|---|
| Stable role and tool instructions | 600 |
| Current workflow step and validated state | 350 |
| Recent conversation | 1,400 |
| Retrieved material and tool results | 900 |
| Reserved output | 300 |
| Template and counting headroom | 546 |
| Total | 4,096 |
Count the fully rendered chat template with the deployed tokenizer, including tool schemas. Before sending, require rendered_input_tokens + reserved_output_tokens <= deployment_limit. Character counts are a poor substitute, especially when a conversation changes language.
Reduce duplicated instructions and oversized tool payloads first. Retrieve the relevant policy section instead of the entire manual. Preserve valid assistant-tool message pairs when trimming history. If required information still does not fit, ask a focused question or route the task; silently dropping the caller’s last correction is not a useful optimization.
Preserve facts when compressing history
Keep a typed record of selected options, explicit confirmations, source turn IDs and revisions. A correction should invalidate dependent state: changing an appointment date invalidates the old availability result and any confirmation for that result. Let a summary preserve conversational context, while server checks decide whether an action is permitted.
If the prompt carries every possible branch on every turn, scope instructions to the current workflow step. If its rules contradict each other, resolve those conflicts before adding context. A larger context window does not resolve either problem.
Common questions
Does prefix caching make every generated token faster?
No. Prefix caching reuses previously computed prompt state, reducing repeated prefill work. It does not directly remove the work of decoding new output tokens.
Should every voice request use the model’s maximum context window?
No. The supported window is a limit, not a target. Budget context for the current task and reserve output space, then measure capacity and quality at realistic history lengths.
Can a conversation summary replace structured application state?
No. Keep confirmed facts, revisions, tool results and permissions in validated application state. A generated summary can help recover conversational context but should not authorize an action.
Sources and API references
References reviewed on . Measurement dates and experimental limits are stated in the article. Check current API and runtime documentation when integrating.