Articles

Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.

New to the terminology? Browse the AI inference glossary.

·3 min read

DeepSeek V4 Pro on AgentX: B200 vs B300 and the KV Cache Working Set

50% more HBM squeezes out extra throughput, and the per-point telemetry shows exactly where it comes from

agentxagenticbenchmarkinferencedeepseekb200b300h200nvidia
·2 min read

GLM 5.3 on AgentX: NVIDIA Is Up to 5x Cheaper per Token at 150 tok/s/user

At this operating point, free AMD silicon would still not close the gap

agentxagenticbenchmarkinferenceglm5b200b300mi355xsglangnvidiaamd
·3 min read

MiniMax M3 on AgentX: Why B200 and B300 Beat Their Rack-Scale GB200 NVL72 & GB300 NVL72 Counterparts

The Dynamo router becomes the bottleneck, no submission runs context parallelism, and AMD leaves KV offload on the table

agentxagenticbenchmarkinferenceminimaxb200b300gb200gb300nvl72dynamoamdrocm
·2 min read

MiniMax M3 on AgentX: B300 TRT-LLM TP2 Owns the Crown

NVIDIA sweeps the 432B model, and the missing DP-attention points explain why cache locality became a routing constraint

agentxagenticbenchmarkinferenceminimaxb300b200gb200trtllmvllmnvidia
·2 min read

Qwen3.5 397B on AgentX: B300 FP4 Delivers 12x the Performance per Dollar of H100

What four years of hardware and a 4-bit format buy on a long-context agentic workload

agentxagenticbenchmarkinferenceqwenb300h100h200fp4fp8sglangnvidia
·23 min read

Kimi K3: The Manos, The Mythos, The Legendos

Kimi K3's architecture: compressed memory, attention across depth, latent expert routing, and serving performance

inferencebenchmarkgpukimivllmnvidiab200b300dynamo
·29 min read

DeepSeekV4 1.6T Day 0 to Day 43 Performance Over Time — Huawei, GB300 NVL72, MI355X, B200

Day 0 Inference Performance, InferenceX, 100x performance improvement in 26 Days, Cost per Million Tokens, Huawei 950DT Inference Trace Analysis

benchmarkgpuinferencedeepseeknvidiaamdhuaweigb300b300b200mi355xh200sglangvllmtrtllmcann