Articles

Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.

New to the terminology? Browse the AI inference glossary.

·2 min read

DeepSeek V4 Pro on AgentX: GB200 vs GB300 Rack-Scale Disaggregation

Both lean on PD disagg, GB300 adds DEP32 wide-EP decode, and the gap shows up in first-token latency rather than token rate

agentxagenticbenchmarkinferencedeepseekgb200gb300nvl72disaggwide-epnvidia
·3 min read

MiniMax M3 on AgentX: Why B200 and B300 Beat Their Rack-Scale GB200 NVL72 & GB300 NVL72 Counterparts

The Dynamo router becomes the bottleneck, no submission runs context parallelism, and AMD leaves KV offload on the table

agentxagenticbenchmarkinferenceminimaxb200b300gb200gb300nvl72dynamoamdrocm
·2 min read

MiniMax M3 on AgentX: B300 TRT-LLM TP2 Owns the Crown

NVIDIA sweeps the 432B model, and the missing DP-attention points explain why cache locality became a routing constraint

agentxagenticbenchmarkinferenceminimaxb300b200gb200trtllmvllmnvidia
·20 min read

Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis

Rubin LUT Based Tensor Core, Feynman, Rack Scale, Perf Per MegaWatt, Perf Per Dollar, Software Improvements, Public Rubin Software, PyTorch, vLLM, OpenAI Triton

benchmarkgpuinferencenvidiarubingb200gb300deepseektrtllmdynamo
·10 min read

GB300 NVL72 vs GB200 NVL72 Inference Performance & Perf per Dollar - on DeepSeek-V4-Pro 1.6T: Up to 2.83x Throughput

DSv4-Pro FP4 8K/1K, Dynamo+vLLM, disaggregated on both racks. GB300's 50% extra HBM (288 vs 192 GB/GPU) unlocks a wider prefill+decode recipe GB200 can't fit — lifting middle-of-curve perf/$ by 2.31x despite a 20% per-GPU TCO premium.

benchmarkgpuinferencedeepseeknvidiagb300gb200nvl72vllmdynamowide-epdisagg
·8 min read

GB200 NVL72 vs B200 on DeepSeek R1 670B: Up to 4.4x Throughput per GPU at 125 tok/s/user

DeepSeek R1 FP4 1k/1k. NVL72's 72-GPU NVLink scale-up fabric lets decode run wide EP up to EP=32, where B200's 8-GPU NVLink island caps out at EP=8 over RoCEv2

benchmarkgpuinferencedeepseeknvidiagb200b200nvl72trtllmdynamowide-epdisagg
·6 min read

GB200 NVL72 vs B200 on Kimi K2.5: 3.1x from Wide EP vLLM

Rack scale NVLink on NVL72 lets Dynamo vLLM run Kimi K2.5 wide EP up to Decode EP 16, taking peak throughput from 4,021 to 12,587 tok/s/GPU on 8k/1k NVFP4

benchmarkgpuinferencekiminvidiagb200b200vllmnvl72wide-ep