OpenAI's Latest In House Chip verus Rubin NVL72New

Compare Jalapeño (Teacup) with Vera Rubin (July) NVL72 on DeepSeek R1 at 8K / 1K.

AgentX / live results

Compare Realistic Agentic Inference Perf

Long Context Multi Turn Inference Performance. Compare Across OpenAI Jalapeño, MI355X, GB300 NVL72, GB200 NVL72, B200, H200, H100, RTX Pro, and soon TPUv7/v8 & Rubin NVL72 & MI455X UALoE72

Every Result Is Transparently done through Public GitHub Actions Automation

Every data point on the dashboard is produced by a public GitHub Actions workflow run. The recipe lives in the repo, the run executes on the actual target hardware, and the full logs and artifacts are publicly viewable. Click any point on a chart to jump straight to the run that produced it. All reproducible, auditable, and open source.

1,000+ new benchmark datapoints added per week on average. Browse every new model, chip, framework, and configuration as it lands.

Public Actions runs
Every benchmark executes on GitHub Actions with full logs visible while the run is in progress.
Open recipes
Every model, framework, precision, and parallelism setting is committed to the public repo as a shell script.
Weekly DB snapshots
The full benchmark database is published as a public GitHub Release every week so the historical dataset stays auditable.