DeepSeek V4 Pro 0813 1.6T
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
Hyperscaler cost ↓ Lower is better Source: InferenceX & SemiAnalysis Market July 2026 Pricing Surveys & AI Cloud TCO Model
| Model · Scenario | B200 | MI355X | B300 | GB200 NVL72 | GB300 NVL72 |
|---|---|---|---|---|---|
DeepSeek V4 Pro 0813 1.6TLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | llm-d vLLM · FP4 | Dynamo vLLM · FP4 | |||
Kimi K3 2.8TLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | vLLM · FP4 | vLLM · FP4 | Dynamo vLLM · FP4 | ||
MiniMax M3 428BLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | — no exact @50 result | vLLM · FP4 | — no exact @50 result | — no exact @50 result | — no exact @50 result |
GLM5.2/GLM5.3 744BLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | — no exact @50 result | — no exact @50 result | — no exact @50 result | — no exact @50 result | — no exact @50 result |
Qwen3.5 397BLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | — no exact @50 result | SGLang · FP4 | SGLang · FP8 | — no exact @50 result | — no exact @50 result |
Qwen3.8 Flash Next 176BLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | — no data for this scenario | — no data for this scenario | — no exact @50 result | — no data for this scenario | — no data for this scenario |
DeepSeek V4 Pro 0813 1.6T8K/1K | TRTLLM · FP4 | ATOM¹ · FP4 | TRTLLM · FP4 | Dynamo vLLM · FP4 | Dynamo SGLang · FP4 · STP |
Qwen3.5 397B8K/1K | TRTLLM · FP4 | SGLang · FP4 | Dynamo SGLang · FP8 · STP | ||
DeepSeek R1 0528 671BMaintenanceModel is no longer actively benchmarked.8K/1K | TRTLLM · FP4 | ATOM¹ · FP4 | Dynamo TRTLLM · FP4 | Dynamo SGLang · FP8 · STP | Dynamo TRTLLM · FP8 · STP |
Kimi K2.5/2.6/2.7-Code 1TDeprecatedModel is no longer actively benchmarked.8K/1K | Dynamo TRTLLM · FP4 · STP | ATOM¹ · FP4 · STP | vLLM · FP4 · STP | Dynamo TRTLLM · FP4 · STP | Dynamo TRTLLM · FP4 · STP |
GLM5/5.1 744BDeprecatedModel is no longer actively benchmarked.8K/1K | SGLang · FP4 | SGLang · FP8 | SGLang · FP4 | Dynamo TRTLLM · FP4 · STP | Dynamo TRTLLM · FP4 · STP |
gpt-oss 120BDeprecatedModel is no longer actively benchmarked.8K/1K | TRTLLM · FP4 · STP | ATOM¹ · FP4 · STP | — no data for this scenario | — no exact @50 result | — no data for this scenario |
MiniMax M2.5/2.7 230BDeprecatedModel is no longer actively benchmarked.8K/1K | vLLM · FP4 · STP | ATOM¹ · FP4 · STP | vLLM · FP4 · STP | Dynamo vLLM · FP4 · STP | Dynamo vLLM · FP4 · STP |
Llama 3.3 70B InstructDeprecatedModel is no longer actively benchmarked.8K/1K | TRTLLM · FP4 · STP | vLLM · FP4 · STP | — no data for this scenario | — no data for this scenario | — no data for this scenario |
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
8K/1K
8K/1K
8K/1K
8K/1K
8K/1K
8K/1K
8K/1K
8K/1K
Current cost and change versus the latest validated platform result 30–60 days earlier.
Platforms without a valid 30-day comparison show current cost only.
If a chip does not have FP4 spec decoding available, the next best available configuration is used.