Original Reddit post

NVIDIA’s new AgentX results replay production-style coding-agent sessions with long context, KV-cache reuse, tool gaps, and dynamic concurrency. Its Vera Rubin preview result claims up to 30× higher throughput per megawatt than GB300 at a matched interactivity target; NVIDIA says the result is pending SemiAnalysis review. This is more representative than fixed 8K/1K serving, but the numerator still stops at tokens. A production benchmark should also report:

Originally posted by u/Crescitaly on r/ArtificialInteligence