InferenceXbySemiAnalysis logo
HomeAgentXNEWOverviewDashboardComparisonsArticlesAbout
Star1,583中文

Articles

Articles on agentic inference, AgentX results, chip performance, and ML infrastructure.

New to the terminology? Browse the AI inference glossary.

Allagenticagentsagentxamdannouncementasicb200b300benchmarkcanndeepseekdisaggdynamofp4fp8gb200gb300glm5gpuh100h200huaweiinferencekimilatencymi355xminimaxnvfp4nvidianvl72openaiqwenrocmrubinsglangthroughputtilerttrtllmvllmwide-ep
August 24, 2026·64 min read

AgentX - InferenceXv3: Does the CUDA Moat Hold Up in Agentic Inferencing?

$3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355X, B200

agentxagenticbenchmarkinferencegpunvidiaamdannouncement
February 16, 2026·47 min read

InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper - Formerly InferenceMAX

GB300 NVL72, MI355X, B200, H100, Disaggregated Serving, Wide Expert Parallelism, Large Mixture of Experts, SGLang, vLLM, TRTLLM

benchmarkgpuinferenceannouncement
October 9, 2025·37 min read

InferenceMAX: Open Source Inference Benchmarking

NVIDIA GB200 NVL72, AMD MI355X, Throughput Token per GPU, Latency Tok/s/user, Perf per Dollar, Cost per Million Tokens, Tokens per Provisioned Megawatt, DeepSeek R1 670B, GPTOSS 120B, Llama3 70B

benchmarkgpuinferenceannouncement
SemiAnalysis logo

Continuous open-source agentic inference benchmarking. Real-world, reproducible, auditable performance data trusted by trillion dollar AI infrastructure operators like OpenAI, Meta, Oracle, Microsoft, etc.

SemiAnalysisMain SiteNewsletterAbout
LegalLand AcknowledgementPrivacy PolicyCookie Policy
ContributeBenchmarksAgentX HarnessVisualization
More
SupportersAgentXTelemetryArticlesAPI ReferenceChip ReliabilityPerformance per DollarModel ArchitecturesAI Inference GlossaryChip Specs & PricingGPU RankingsModel on GPU Results

If this data helps your work, consider starring us on GitHub or sharing with your network.

© 2026 semianalysis.com. All rights reserved.