For developers
Build the conversation.
Understand the stack.
Measured serving trade-offs and the infrastructure behind a production call. Explore Gemma, FP4, vLLM and context management, then inspect the data and run your own experiments.
Voice infrastructure
-
Making Gemma 4 26B A4B fast enough for a voice conversation
A Gemma voice-serving case study: 1,520 measured requests, concurrency trade-offs, connection reuse and a downloadable analysis kit.
-
How FP4 can help voice inference—and what it cannot promise
Understand NVFP4 with Gemma 4 26B A4B: weight memory, hardware kernels, FP8 KV cache and the experiment needed to prove a latency improvement.
-
Tune vLLM for voice latency, not just tokens per second
Use our archived concurrency sweep to understand vLLM batching, tail latency, prefill budgets, warmup and goodput for voice workloads.
-
Manage voice-agent context before buying more inference capacity
Measured shared-prefix results, KV capacity arithmetic and a practical context budget for voice agents using Gemma and vLLM.
Looking for endpoint fields and response schemas? Open the API documentation. For release decisions and measurements, read the engineering journal.