Articles

Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.

New to the terminology? Browse the AI inference glossary.

·3 min read

DeepSeek V4 Pro on AgentX: MI355X vs B200, and the August 21 Flip

AMD matched B200 vLLM on performance per dollar for end-to-end latency, then upstream vLLM work moved the line

agentxagenticbenchmarkinferencedeepseekmi355xb200amdnvidia
·2 min read

GLM 5.3 on AgentX: MI355X ATOM Beats GB300 NVL72 on Part of the Curve

Where AMD’s vendor engine wins on performance per dollar, and what E2E Normalized Interactivity actually measures

agentxagenticbenchmarkinferenceglm5mi355xgb300nvl72trtllmsglangamdnvidia
·2 min read

GLM 5.3 on AgentX: NVIDIA Is Up to 5x Cheaper per Token at 150 tok/s/user

At this operating point, free AMD silicon would still not close the gap

agentxagenticbenchmarkinferenceglm5b200b300mi355xsglangnvidiaamd
·2 min read

Kimi K3 on AgentX: MI355X ATOM Beats GB300 NVL72 on Part of the Curve

AMD’s vendor engine wins a real slice of the performance per dollar frontier, while Hopper struggles to serve K3 at all

agentxagenticbenchmarkinferencekimimi355xgb300nvl72h200amdnvidia
·3 min read

MiniMax M3 on AgentX: Why B200 and B300 Beat Their Rack-Scale GB200 NVL72 & GB300 NVL72 Counterparts

The Dynamo router becomes the bottleneck, no submission runs context parallelism, and AMD leaves KV offload on the table

agentxagenticbenchmarkinferenceminimaxb200b300gb200gb300nvl72dynamoamdrocm
·2 min read

MI355X versus GB300 NVL72 Inference Performance: 20x Gap on Qwen3.5 SGLang

GatedDeltaNet, a 262k native context, and no AMD competition at all on the same engine

agentxagenticbenchmarkinferenceqwensglangtrtllmnvidiaamd
·64 min read

AgentX - InferenceXv3: Does the CUDA Moat Hold Up in Agentic Inferencing?

$3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355X, B200

agentxagenticbenchmarkinferencegpunvidiaamdannouncement
·29 min read

DeepSeekV4 1.6T Day 0 to Day 43 Performance Over Time — Huawei, GB300 NVL72, MI355X, B200

Day 0 Inference Performance, InferenceX, 100x performance improvement in 26 Days, Cost per Million Tokens, Huawei 950DT Inference Trace Analysis

benchmarkgpuinferencedeepseeknvidiaamdhuaweigb300b300b200mi355xh200sglangvllmtrtllmcann
·13 min read

MI355X DeepSeek-V4-Pro on SGLang: 110.5x Throughput per GPU in 26 Days

The amd/deepseek_v4 side branch shipped TileLang attention indexer, Triton sparse MLA, fused RoPE/Hadamard, FlyDSL MoE, and FP4 weights across 31 performance optimizations PRs — lifting first-light 20 tok/s/GPU at 2.4 tok/s/user into 2,256 tok/s/GPU at 9.4 tok/s/user on 8K/1K, with both throughput and interactivity climbing together

benchmarkgpuinferencedeepseekamdmi355xsglangrocmfp4
·7 min read

AMD MI355X GLM-5 Inference: Up to 40% Cheaper per Million Tokens than B200 on SGLang FP8

14 weeks after GLM-5 launched, AMD landed both MTP and non-MTP SGLang FP8 recipes on MI355X — fused MLA + FP8 KV cache via TileLang flips the single-node FP8 cost curve in AMD favor across most of the performance Pareto

benchmarkgpuinferenceglm5amdnvidiami355xb200sglangrocm
·6 min read

AMD MI355X Qwen3.5 397B-A17B Inference: Up to 19x Throughput per GPU in 3 Months on SGLang FP8

From v0.5.8 (Feb) → v0.5.10rc0 (Apr) → v0.5.12 (May), three AITER kernel landings on MI355X plus a TP=8 → TP=2/TP=4 retune push Qwen3.5 8k/1k peak from 1.3k to 6.4k tok/s/GPU and extend the curve out to 75 tok/s/user

benchmarkgpuinferenceqwenamdmi355xsglangrocm
·6 min read

AMD MI355X Kimi K2.5 Inference: 7.7x Throughput, Up To 15x Interactivity in 25 Days on vLLM

vLLM PR #35850 Fixed AITER MLA Dispatch on MI355X CDNA4, Unlocking Kimi K2.5 Inference Performance at TP=8, Shipped in vLLM 0.18

benchmarkgpuinferencekimiamdvllmrocmmi355x