← Tutte le news
ARTICOLO
22 settembre 2026

Vera Rubin NVL72 MLPerf preview — solo numeri NVIDIA

NVIDIA blog 2026-09-16. Preview MLPerf v6.1: fino a ~3.7× / ~2.5× NVIDIA-claimed vs GB300. Evitare #1 assoluto senza MLCommons.

On , NVIDIA published a first preview MLPerf Inference v6.1 submission for Vera Rubin NVL72. Headline figures are NVIDIA-claimed: up to ~3.7× throughput vs GB300 NVL72 on Qwen3-VL (vLLM + Dynamo) and up to ~2.5× on DeepSeek-R1 (TensorRT-LLM). MLCommons marks the configuration in a Preview category — early structured results, not proof of broad commercial availability or stable rental economics.

Editorial rule: do not crown absolute “#1” without MLCommons language and entry IDs. Quote ratios as NVIDIA-stated preview comparisons to like-sized GB300 NVL72, and note scenario variance (interactive gains can differ from offline/server). Power, TCO, and fleet reliability are out of scope for this draft.

MSP takeaway: useful for capacity planning conversations, but gate POCs on your own SLOs and software stack maturity — preview silicon ≠ production SLA.

FAQ

Sono vincitori ufficiali MLCommons?

Sono submission preview NVIDIA in MLPerf Inference v6.1. Riportare i rapporti NVIDIA-claimed e lo status Preview; evitare #1 assoluto senza framing MLCommons.

Cosa devono ancora validare gli MSP?

Stack software, power/cooling e SLO di produzione — il throughput preview non è una garanzia managed-service.

Sources

Draft only — do not publish without editorial review.