GLM 5.3 on AgentX: MI355X ATOM Beats GB300 NVL72 on Part of the Curve
Where AMD’s vendor engine wins on performance per dollar, and what E2E Normalized Interactivity actually measures
Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.
New to the terminology? Browse the AI inference glossary.
Where AMD’s vendor engine wins on performance per dollar, and what E2E Normalized Interactivity actually measures
NVIDIA sweeps the 432B model, and the missing DP-attention points explain why cache locality became a routing constraint
GatedDeltaNet, a 262k native context, and no AMD competition at all on the same engine
Rubin LUT Based Tensor Core, Feynman, Rack Scale, Perf Per MegaWatt, Perf Per Dollar, Software Improvements, Public Rubin Software, PyTorch, vLLM, OpenAI Triton
Day 0 Inference Performance, InferenceX, 100x performance improvement in 26 Days, Cost per Million Tokens, Huawei 950DT Inference Trace Analysis
DeepSeek R1 FP4 1k/1k. NVL72's 72-GPU NVLink scale-up fabric lets decode run wide EP up to EP=32, where B200's 8-GPU NVLink island caps out at EP=8 over RoCEv2