·2 min read
Qwen3.5 397B on AgentX: B300 FP4 Delivers 12x the Performance per Dollar of H100
What four years of hardware and a 4-bit format buy on a long-context agentic workload
agentxagenticbenchmarkinferenceqwenb300h100h200fp4fp8sglangnvidia
Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.
New to the terminology? Browse the AI inference glossary.
What four years of hardware and a 4-bit format buy on a long-context agentic workload
GatedDeltaNet, a 262k native context, and no AMD competition at all on the same engine
From v0.5.8 (Feb) → v0.5.10rc0 (Apr) → v0.5.12 (May), three AITER kernel landings on MI355X plus a TP=8 → TP=2/TP=4 retune push Qwen3.5 8k/1k peak from 1.3k to 6.4k tok/s/GPU and extend the curve out to 75 tok/s/user