# Voice infrastructure field notes — Mirai Minds

Published September 17, 2026. Measurements were recorded in July and September 2026; they were not rerun for publication.

## What is in the kit

- `observations.json`: allowlisted numeric observations, timing boundaries and limitations.
- `summary.csv`: pooled results computed from those observations.
- `analyze.py`: Python 3 standard-library reanalysis; no GPU or network access.
- `concurrency.png` and `concurrency.svg`: plots of the archived concurrency sweep.
- `plot.py`: regenerate the charts with matplotlib installed.
- `prompt-workshop.md`: fictional before/after prompt, state contract and evaluation method.
- `eval-cases.json`: 12 authored workflow scenarios; no model results implied.
- `experiment-worksheet.md`: a template for a new matched serving experiment.

Run `python3 analyze.py` from the extracted directory. It reads the adjacent observations file, rewrites summary.csv and prints the pooled statistics. Percentiles use linear interpolation at `(n - 1) * p` over all saved request timings for a condition. Do not average the individual runs' p95 values.

For charts, install matplotlib in your own Python environment and run `python3 plot.py`.

## Provenance and scope

Deployment documentation maps the `voice-agent` alias to `nvidia/Gemma-4-26B-A4B-NVFP4`. The archived benchmark results record the alias, not an immutable checkpoint revision or full server manifest. The current vLLM 0.26.0 pin cannot establish the precise runtime used by every earlier run.

The July 28 concurrency sweep ran on RTX PRO 6000 Blackwell (96 GB). It covers six concurrency levels, five repetitions per level, one discarded warmup wave and two measured waves per repetition: 1,520 measured requests total. Each request opened a fresh TCP connection. The timer ran from before connection establishment to the first nonempty content delta. It excluded ASR, TTS, playback, reasoning output and tool completion.

A shared 3,403-character system prompt described a short Hindi/English voice-support conversation. The user reference varied. The prompt-token sample recorded in every shared-prefix run was 758. Generation temperature was 0.4, with max_tokens 80. Original domain-specific prompt text and response text are not redistributed. This means exact inference replay is unavailable; the supplied data supports analysis replay.

The July 27 prefix comparison ran on RTX 5090 (32 GB), with three repetitions of each condition at concurrency 16: 96 measured requests each. A unique nonce before the system prompt changed the recorded prompt-token sample to 782. The two conditions ran sequentially. Cache-hit telemetry was not recorded. This is not a fully isolated caching ablation.

The September 8 connection probe contains 12 serial requests seven seconds apart, one new connection, HTTP 200 for every request, and a 120-second keep-alive timeout. Its first-useful-output boundary differs from July's content-only harness. Exact checkpoint and hardware revisions were not archived with this probe. Do not combine these samples with the July distributions.

## Instrumentation limitations that must travel with the numbers

1. The legacy July harness preferred server usage completion_tokens but fell back to counting nonempty content chunks when usage was absent. It did not archive which counter each request used. Published throughput is an estimate of recorded output per measured second, not an audited token rate.
2. The legacy SSE parser did not preserve incomplete lines between reads. It could skip fragmented events. Raw streams were not archived for re-parsing. These are the harness's observations, not independently verified service timings.
3. Saved request timings were rounded to 0.1 ms; measured wave durations were rounded to 0.001 s. Recomputed rates can differ slightly from the original unrounded calculation.
4. Warmup and pauses are excluded from throughput. Short synchronized bursts with a warmed prefix do not establish sustained arrival-rate capacity, full-context capacity or concurrent phone-call capacity.
5. No task-quality scores or matched BF16-versus-NVFP4 latency experiment are included. Zero harness errors is not proof that the answers were correct.
6. GPU-memory arithmetic and proposed prompt budgets in the guides are labelled calculations or examples. They are not additional measurements.

## Attribution

When sharing an observation, credit Mirai Minds and include the measurement date, workload, timing boundary and limitations above. The authored workshop and scripts accompany the guides at https://voice.miraiminds.co/guides. External model and software documentation retain their own terms.
