AgentX / live results

Compare Realistic Agentic Inference Perf

Long Context Multi Turn Inference Performance. Compare Across OpenAI Jalapeño, MI355X, GB300 NVL72, GB200 NVL72, B200, H200, H100, RTX Pro, and soon TPUv7/v8 & Rubin NVL72 & MI455X UALoE72

Comparison catalog

AgentX and 8K→1K results

305 head-to-head inference benchmark comparisons across DeepSeekv4 Pro 0813 1.6T, DeepSeek R1, Kimi K3 2.8T, Kimi K2.5/K2.6/K2.7-Code 1T, GLM 5/5.1, GLM 5.3 744B, MiniMax M3 428B, MiniMax M2.5/M2.7, Qwen 3.8 Flash Next 176B-A6B, Qwen 3.5 397B-A17B, gpt-oss 120B, and Llama 3.3 70B. Models with AgentX data open long-context, multi-turn trace replay results. Models not yet covered by AgentX open the controlled 8K→1K workload. Each card identifies its scenario.

DeepSeek R1

36 chip pairs with benchmark data on DeepSeek R1.

MiniMax M3 428B

36 chip pairs with benchmark data on MiniMax M3 428B.

MiniMax M2.5/M2.7

36 chip pairs with benchmark data on MiniMax M2.5/M2.7.

Qwen 3.8 Flash Next 176B-A6B

3 chip pairs with benchmark data on Qwen 3.8 Flash Next 176B-A6B.

Qwen 3.5 397B-A17B

45 chip pairs with benchmark data on Qwen 3.5 397B-A17B.