InferenceXbySemiAnalysis logo
HomeAgentXNEWOverviewDashboardComparisonsArticlesAbout
Star1,583中文

Articles

Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.

New to the terminology? Browse the AI inference glossary.

Allagenticagentsagentxamdannouncementasicb200b300benchmarkcanndeepseekdisaggdynamofp4fp8gb200gb300glm5gpuh100h200huaweiinferencekimilatencymi355xminimaxnvfp4nvidianvl72openaiqwenrocmrubinsglangthroughputtilerttrtllmvllmwide-ep
August 25, 2026·2 min read

DeepSeek V4 Pro on AgentX: GB200 vs GB300 Rack-Scale Disaggregation

Both lean on PD disagg, GB300 adds DEP32 wide-EP decode, and the gap shows up in first-token latency rather than token rate

agentxagenticbenchmarkinferencedeepseekgb200gb300nvl72disaggwide-epnvidia
May 27, 2026·10 min read

GB300 NVL72 vs GB200 NVL72 Inference Performance & Perf per Dollar - on DeepSeek-V4-Pro 1.6T: Up to 2.83x Throughput

DSv4-Pro FP4 8K/1K, Dynamo+vLLM, disaggregated on both racks. GB300's 50% extra HBM (288 vs 192 GB/GPU) unlocks a wider prefill+decode recipe GB200 can't fit — lifting middle-of-curve perf/$ by 2.31x despite a 20% per-GPU TCO premium.

benchmarkgpuinferencedeepseeknvidiagb300gb200nvl72vllmdynamowide-epdisagg
May 23, 2026·8 min read

GB200 NVL72 vs B200 on DeepSeek R1 670B: Up to 4.4x Throughput per GPU at 125 tok/s/user

DeepSeek R1 FP4 1k/1k. NVL72's 72-GPU NVLink scale-up fabric lets decode run wide EP up to EP=32, where B200's 8-GPU NVLink island caps out at EP=8 over RoCEv2

benchmarkgpuinferencedeepseeknvidiagb200b200nvl72trtllmdynamowide-epdisagg
SemiAnalysis logo

Continuous open-source agentic inference benchmarking. Real-world, reproducible, auditable performance data trusted by trillion dollar AI infrastructure operators like OpenAI, Meta, Oracle, Microsoft, etc.

SemiAnalysisMain SiteNewsletterAbout
LegalLand AcknowledgementPrivacy PolicyCookie Policy
ContributeBenchmarksAgentX HarnessVisualization
More
SupportersAgentXTelemetryArticlesAPI ReferenceChip ReliabilityPerformance per DollarModel ArchitecturesAI Inference GlossaryChip Specs & PricingGPU RankingsModel on GPU Results

If this data helps your work, consider starring us on GitHub or sharing with your network.

© 2026 semianalysis.com. All rights reserved.