DeepSeek V4 Pro on AgentX: B200 vs B300 and the KV Cache Working Set
50% more HBM squeezes out extra throughput, and the per-point telemetry shows exactly where it comes from
Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.
New to the terminology? Browse the AI inference glossary.
50% more HBM squeezes out extra throughput, and the per-point telemetry shows exactly where it comes from
At this operating point, free AMD silicon would still not close the gap
The Dynamo router becomes the bottleneck, no submission runs context parallelism, and AMD leaves KV offload on the table
NVIDIA sweeps the 432B model, and the missing DP-attention points explain why cache locality became a routing constraint
What four years of hardware and a 4-bit format buy on a long-context agentic workload
Kimi K3's architecture: compressed memory, attention across depth, latent expert routing, and serving performance
Day 0 Inference Performance, InferenceX, 100x performance improvement in 26 Days, Cost per Million Tokens, Huawei 950DT Inference Trace Analysis