GB300 NVL72 FP4: MTP vs Off Speculative Decoding
Speculative decoding comparison of MTP versus Off on GB300 NVL72 FP4 (NVIDIA Blackwell) running Qwen 3.5 397B-A17B. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.
MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.
Near the low end of the 46–278 tok/s/user interactivity band, at 104 tok/s/user on Qwen 3.5 397B-A17B (GB300 NVL72 FP4): MTP runs 34137 tok/s/chip at $0.02/M tokens, Off runs 12387 at $0.05/M. MTP is 176% cheaper per token; MTP delivers 176% more tok/s/chip. Gains from speculative decoding vary by workload; short-output prompts tend to benefit less.
At 162 tok/s/user on Qwen 3.5 397B-A17B (GB300 NVL72 FP4), MTP delivers 18210 tok/s/chip at $0.04 per million tokens; Off delivers 5253 tok/s/chip at $0.12. MTP is 247% cheaper per token; MTP delivers 247% more tok/s/chip. Speculative decoding accepts draft tokens to reduce per-token latency — gains vary by workload and prompt distribution.
MTP posts 12786 tok/s/chip for $0.05 per million tokens at 220 tok/s/user on Qwen 3.5 397B-A17B (GB300 NVL72 FP4); Off posts 1895 tok/s/chip for $0.34. MTP is 575% cheaper per token; MTP delivers 575% more tok/s/chip. Draft-token acceptance rates determine whether speculative decoding helps or hurts at a given concurrency level. (Numbers reflect this URL's pinned 8k/1k · fp4 workload — changing sequence or model updates both the table and chart; the table stays pinned to this page's precision, so precision toggles in the controls affect the chart only.)

| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | MTP:34136.6Off:12387.4 | MTP:18209.8Off:5253.1 | MTP:12786.4Off:1895.2 |
| Cost ($/M tok) | MTP:$0.019Off:$0.052 | MTP:$0.035Off:$0.122 | MTP:$0.050Off:$0.339 |
| tok/s/MW | MTP:16102183Off:5843118 | MTP:8589552Off:2477888 | MTP:6031338Off:893960 |
| Concurrency | MTP:~1171Off:~307 | MTP:~479Off:~118 | MTP:~166Off:~32 |