DeepSeek V4 Pro on AgentX: B200 vs B300 and the KV Cache Working Set
50% more HBM squeezes out extra throughput, and the per-point telemetry shows exactly where it comes from
Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.
New to the terminology? Browse the AI inference glossary.
50% more HBM squeezes out extra throughput, and the per-point telemetry shows exactly where it comes from
Both lean on PD disagg, GB300 adds DEP32 wide-EP decode, and the gap shows up in first-token latency rather than token rate
AMD matched B200 vLLM on performance per dollar for end-to-end latency, then upstream vLLM work moved the line
Where AMD’s vendor engine wins on performance per dollar, and what E2E Normalized Interactivity actually measures
At this operating point, free AMD silicon would still not close the gap
AMD’s vendor engine wins a real slice of the performance per dollar frontier, while Hopper struggles to serve K3 at all
The Dynamo router becomes the bottleneck, no submission runs context parallelism, and AMD leaves KV offload on the table
NVIDIA sweeps the 432B model, and the missing DP-attention points explain why cache locality became a routing constraint
What four years of hardware and a 4-bit format buy on a long-context agentic workload
GatedDeltaNet, a 262k native context, and no AMD competition at all on the same engine
$3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355X, B200
Multi-turn sessions, long contexts, and near-total prefix reuse make agentic inference a systems problem, and change what a benchmark has to measure
Can TileRT software on NVIDIA GPUs compete with Cerebras, Groq LPU, and SambaNova? Batch size 1, disaggregated engine, high-throughput prefill engine, high-interactivity decode engine