WORLDTECH NEWS Global technology intelligence.Contact
← Back to WORLDTECH

Red Hat delivers peak performance on Kubernetes and CPUs in MLPerf Inference v6.1

Detailed view of hands installing a CPU onto a motherboard inside a computer setup.
Illustrative photo.Photo by Athena Sandrini on Pexels

What happened

The full specifications of the Supermicro server can be found here .This was the only submission in the closed datacenter division running on Kubernetes (software that runs and manages applications across many servers) infrastructure. Notably, 22 of 32 datacenter submitters in this round used vLLM somewhere in their stack, confirming its position as the industry-standard open source inference engine. GPT-OSS-120B on GB200 NVL4 with Red Hat OpenShift (reasoning model) GPT-OSS-120B is a 117-billion-parameter mixture-of-experts model for reasoning, agentic (AI that carries out multi-step tasks rather than answering one question) workflows, and code generation, with highly variable input lengths and strict server-scenario latency (the delay between asking for something and getting it) requirements.

Key facts

  • The full specifications of the Supermicro server can be — found: here .This was the only submission in the closed datacenter division running on Kubernetes infrastructure

Sources & evidence