Voice infrastructure

How FP4 can help voice inference—and what it cannot promise

FP4 can reduce weight traffic and let supported hardware execute lower-precision operations. That can help LLM serving, but the latency gain depends on kernels, batch size, context and quality requirements. Our Gemma checkpoint uses an experts-only NVFP4 recipe; an FP4 label does not mean every tensor, or the KV cache, uses four bits.

By Mirai Minds · Published

Understand the two opportunities

A smaller weight representation can reduce the bytes moved while computing a response. Supported low-precision matrix kernels can also change compute time. These are possible routes to faster serving; neither predicts a particular caller’s wait by itself.

For a short voice reply, ask which part of the trace dominates. A queue, a long prompt prefill or a slow connection may outweigh the time spent producing the answer. Reducing weight precision does not remove those waits. Under higher load, any extra capacity may instead be spent on more simultaneous work.

The memory arithmetic behind NVFP4

NVFP4 represents values with four bits, plus an eight-bit scale shared by each block of 16 values and a second-level tensor scale. That is 4.5 bits per value before tensor-level overhead. See NVIDIA’s format description.

For an illustrative tensor containing one billion values, the storage arithmetic is:

Representation Arithmetic Decimal GB
BF16 values 1 billion × 16 ÷ 8 2.0000
FP8 values, excluding scales 1 billion × 8 ÷ 8 1.0000
Four-bit values only 1 billion × 4 ÷ 8 0.5000
NVFP4 values plus block scales 1 billion × 4.5 ÷ 8 0.5625

The last row is about 3.56 times smaller than the BF16 row. This is storage math for a hypothetical tensor. It excludes other tensors, tensor scales, allocator overhead, activations, runtime workspaces and attention state. It is neither a GPU-memory measurement nor a speedup claim.

What our Gemma checkpoint actually quantizes

nvidia/Gemma-4-26B-A4B-NVFP4 uses the nvfp4_experts_only recipe. Its card specifies Blackwell support. Read the checkpoint configuration and selected runtime backend instead of assuming that a filename describes every operation. Our serving configuration separately sets --kv-cache-dtype fp8. The checkpoint documentation is the authority for its quantization recipe.

This distinction matters when planning capacity. Smaller weights can leave room for request state, but context length and concurrency still have to fit. The model’s small active parameter count does not mean only that many parameters need to be resident.

What happens to accuracy?

NVIDIA publishes the following evaluation results for this checkpoint. These are external model-card results, not Mirai voice-agent tests:

Evaluation Full-precision baseline NVFP4
GPQA Diamond 80.30% 79.90%
MMLU Pro 85.00% 84.80%
IFEval 96.60% 96.40%

Source: NVIDIA Gemma evaluation table. Its published setup uses different generation settings from our short-response serving sweep. These scores do not establish parity on spoken Hindi, tool arguments or booking confirmations.

For a voice application, evaluate the actual failure cost: a wrong digit, a missed correction, an invalid tool argument, an unsupported promise, or an answer that continues after the caller asks to stop. Keep the same fixtures when changing precision.

The experiment that would prove a latency gain

Run both precision candidates on hardware that can hold them. Record the checkpoint revisions, tokenizer, runtime, kernels, prompt tokens, output tokens, cache state, generation settings and client location. First compare at matched concurrency; then compare the maximum arrival rate each can sustain within the same latency and quality thresholds.

Alternate candidate order across repetitions to reduce warmup and time-of-day effects. Keep failures, cancellations and malformed tools in the denominator. Report first useful output, inter-token delay, complete-response time, peak memory and task success separately. If the full-precision model does not fit on the chosen card, label that as a deployment constraint; it is not a measured latency result.

Our archived Gemma sweep supplies useful serving observations. It does not supply this ablation. That is the next experiment to run before attributing a percentage improvement to FP4.

Common questions

Does FP4 make an LLM four times faster than BF16?

No fixed speedup follows from the bit width. Memory traffic, scaling overhead, supported kernels, attention, batching and transport all affect observed latency. A matched workload comparison is required.

Is FP4 quantization the same as an FP8 KV cache?

No. Weight and activation quantization changes model arithmetic and storage. The KV cache stores attention state for requests and has a separate precision setting. Our serving configuration uses an FP8 KV cache.

Has Mirai measured the isolated latency benefit of FP4?

The published archive contains NVFP4 deployment observations, but no matched BF16-versus-NVFP4 latency experiment. We do not claim an isolated quantization speedup from those results.

Sources and API references

References reviewed on . Measurement dates and experimental limits are stated in the article. Check current API and runtime documentation when integrating.

Try the sandbox