InferenceXbySemiAnalysis logo
HomeAgentXNEWOverviewDashboardComparisonsArticlesAbout
Star1,583中文

Articles

Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.

New to the terminology? Browse the AI inference glossary.

Allagenticagentsagentxamdannouncementasicb200b300benchmarkcanndeepseekdisaggdynamofp4fp8gb200gb300glm5gpuh100h200huaweiinferencekimilatencymi355xminimaxnvfp4nvidianvl72openaiqwenrocmrubinsglangthroughputtilerttrtllmvllmwide-ep
August 25, 2026·2 min read

Qwen3.5 397B on AgentX: B300 FP4 Delivers 12x the Performance per Dollar of H100

What four years of hardware and a 4-bit format buy on a long-context agentic workload

agentxagenticbenchmarkinferenceqwenb300h100h200fp4fp8sglangnvidia
May 26, 2026·12 min read

B200 NVFP4 vs H100 FP8 on MiniMax-M2.5: Up to 8.2x Better Performance per Dollar with vLLM

vLLM PR #36307 unlocks the trtllm-gen FP8 MoE kernel for MiniMax on B200; combined with NVFP4, perf/$ scales from 4.0x at 22 tok/s/user to 8.2x at 110 on 8K/1K

benchmarkgpuinferenceminimaxnvidiab200h100vllmfp4
SemiAnalysis logo

Continuous open-source agentic inference benchmarking. Real-world, reproducible, auditable performance data trusted by trillion dollar AI infrastructure operators like OpenAI, Meta, Oracle, Microsoft, etc.

SemiAnalysisMain SiteNewsletterAbout
LegalLand AcknowledgementPrivacy PolicyCookie Policy
ContributeBenchmarksAgentX HarnessVisualization
More
SupportersAgentXTelemetryArticlesAPI ReferenceChip ReliabilityPerformance per DollarModel ArchitecturesAI Inference GlossaryChip Specs & PricingGPU RankingsModel on GPU Results

If this data helps your work, consider starring us on GitHub or sharing with your network.

© 2026 semianalysis.com. All rights reserved.