Articles

Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.

New to the terminology? Browse the AI inference glossary.

·2 min read

GLM 5.3 on AgentX: MI355X ATOM Beats GB300 NVL72 on Part of the Curve

Where AMD’s vendor engine wins on performance per dollar, and what E2E Normalized Interactivity actually measures

agentxagenticbenchmarkinferenceglm5mi355xgb300nvl72trtllmsglangamdnvidia
·2 min read

GLM 5.3 on AgentX: NVIDIA Is Up to 5x Cheaper per Token at 150 tok/s/user

At this operating point, free AMD silicon would still not close the gap

agentxagenticbenchmarkinferenceglm5b200b300mi355xsglangnvidiaamd
·17 min read

Ultra-High Interactivity on NVIDIA GPUs? TileRT on InferenceX

Can TileRT software on NVIDIA GPUs compete with Cerebras, Groq LPU, and SambaNova? Batch size 1, disaggregated engine, high-throughput prefill engine, high-interactivity decode engine

benchmarkgpuinferencenvidiab200gb300tilertvllmglm5agentxagentic
·10 min read

B200 NVFP4 vs H200 FP8 on GLM-5: Up to 3.65x Better Performance per Dollar with SGLang MTP

Both SKUs run SGLang EAGLE MTP; the Blackwell generation lifts perf/$ by ~1.2x at the peak and the NVIDIA GLM-5-NVFP4 checkpoint on FlashInfer TRT-LLM sparse MLA stacks another ~2.4–3.0x on 8K/1K

benchmarkgpuinferenceglm5nvidiab200h200sglangfp4
·7 min read

AMD MI355X GLM-5 Inference: Up to 40% Cheaper per Million Tokens than B200 on SGLang FP8

14 weeks after GLM-5 launched, AMD landed both MTP and non-MTP SGLang FP8 recipes on MI355X — fused MLA + FP8 KV cache via TileLang flips the single-node FP8 cost curve in AMD favor across most of the performance Pareto

benchmarkgpuinferenceglm5amdnvidiami355xb200sglangrocm