Red Hat delivers peak performance on Kubernetes and CPUs in MLPerf Inference v6.1
What happened
The full specifications of the Supermicro server can be found here .This was the only submission in the closed datacenter division running on Kubernetes (software that runs and manages applications across many servers) infrastructure. Notably, 22 of 32 datacenter submitters in this round used vLLM somewhere in their stack, confirming its position as the industry-standard open source inference engine. GPT-OSS-120B on GB200 NVL4 with Red Hat OpenShift (reasoning model) GPT-OSS-120B is a 117-billion-parameter mixture-of-experts model for reasoning, agentic (AI that carries out multi-step tasks rather than answering one question) workflows, and code generation, with highly variable input lengths and strict server-scenario latency (the delay between asking for something and getting it) requirements.
Key facts
- The full specifications of the Supermicro server can be — found: here .This was the only submission in the closed datacenter division running on Kubernetes infrastructure
Sources & evidence
- Red Hat Blog Primary / official
Red Hat delivers peak performance on Kubernetes and CPUs in MLPerf Inference v6.1 ↗
https://www.redhat.com/en/blog/red-hat-delivers-peak-performance-kubernetes-cpus-mlperf-inference-v61