Kimi K3 2.8T · Speculative Decoding
GB300 NVL72 FP4: DSpark vs Off Speculative Decoding
Speculative decoding comparison of DSpark versus Off on GB300 NVL72 FP4 (NVIDIA Blackwell) running Kimi K3 2.8T. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.
MTP acceptance-rate comparability
MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.

No interpolated data available for the default workload on this configuration. Use the chart controls below to select a sequence and precision with benchmark data for both configurations.