Original Reddit post

Hey r/ArtificialInteligence , We’ve published two technical write-ups on serving real-time AI models and wanted to share the main findings here. We build the inference engine; NVIDIA develops the model we tested, Nemotron VoiceChat 11B. A request ends. A session runs on a clock. AI inference today is organized as requests: an input arrives, the model runs, the response ends and its resources are freed. If the server is busy, it can wait a moment to batch work or reorder the queue, and the cost is a slightly slower response. A growing class of models works differently. They stay active alongside something outside the GPU (a conversation, a video stream, a robot), take in input continuously and keep their state for the whole session. We call this continuous inference. Full-duplex voice is the clearest example. The model listens while it speaks, so you can interrupt it. Audio keeps arriving for the whole call, its memory of the conversation stays live until you hang up, and every 80 ms it owes the next frame of output. That deadline comes whether the server is ready or not. If the frame is late, you hear a gap. It also has to run during silence: the length of a pause is how the model tells a hesitation from the end of your turn. Unlike a classic voice agent (speech-to-text → LLM → text-to-speech, where the LLM sits idle between turns), there’s no idle time to skip. Why that gets so expensive Servers like vLLM, SGLang, TensorRT-LLM or Triton were built around requests, and several of their assumptions stop holding: The batch can’t wait. At every tick, the server has to run whatever is due with what it has. Overload hits everyone at once. Sessions batched together share the same step, so one slow step makes all of them late. Memory can’t be freed or swapped out mid-call. It’s needed again 80 ms later. The usual metrics hide failures. Tokens per second and average latency can look fine while one caller hears gaps. In NVIDIA’s reference stack, each 160 ms of audio becomes thousands of small GPU operations with the CPU coordinating between them, and the GPU sits idle about 75% of the time. One session fits. With two, 8–15% of audio beats arrive late. So each live conversation pays for a whole H100, most of which is waiting. How to fix it Not with a faster GPU or a different model: same weights, same precision, same outputs. The fix is in how the model runs. Keep the work on the GPU. Our engine runs the model as one GPU program that stays resident. Every beat, the CPU drops in new audio and picks up the output; nothing else goes back and forth. Advance every live session together. All sessions due on the same tick run as one batch, so the model’s weights are read once for the group instead of once per session. Decide everything ahead of time. The work repeats identically every beat, so a compiler fixes the schedule and memory layout before the first call. No runtime scheduler or allocator adding delays. On a single NVIDIA H100 SXM 80 GB, we measured: 56 concurrent sessions, the highest capacity tested, vs 1 for the reference stack. 147.4–147.5 ms p99 per 160 ms beat. Zero missed deadlines across 84,000 measured session-beats, over three runs. A conversation’s cost is GPU time divided by how many conversations share the GPU, so 56 instead of 1 means about 98% less GPU per conversation. These are server-side measurements that exclude network and audio playback, with the same recorded input across sessions and about two minutes of context. If you’re running real-time models, how many sessions per GPU do you get today, and how do you check that each one stays on time? Links Why duplex inference runs on a clock: https://dotwave.ai/technical-notes/duplex-inference/ How we serve 56 sessions on one H100: https://dotwave.ai/technical-notes/wpk-persistent-kernel/ Browser demo: https://dotwave.ai/demo/ Nemotron VoiceChat 11B (full-duplex) API: https://dotwave.ai/models/nemotron-voicechat-11b/ Nemotron 3.5 ASR Streaming 0.6B API: https://dotwave.ai/models/nemotron-asr-streaming-0-6b/ Sign-up includes free usage, no card required: about 33 hours of Nemotron VoiceChat 11B, or about 370 hours of Nemotron 3.5 ASR Streaming. submitted by /u/Ok_boss_labrunz

Originally posted by u/Ok_boss_labrunz on r/ArtificialInteligence