WORLDTECH NEWS Global technology intelligence.Contact
โ† Back to WORLDTECH

Impactful scheduling for GPU clusters

Detailed close-up image of NVIDIA RTX 2080 graphics card showcasing hardware components.
Illustrative photo.Photo by Nana Dua on Pexels

What happened

Next is impact : how often the most valuable workloads are chosen to receive resources. This occurred because researchers found they could not launch debugging workloads with low enough latency (the delay between asking for something and getting it) to tackle problems in real time.

The foundation is availability : how often the hardware is healthy and ready for work. Above this is occupancy : the fraction of available time assigned to a specific workload.

The capstone of the pyramid is utilization : the fraction of GPU (a chip built for many calculations at once, used for graphics and AI) capacity used over the lifetime of a workload. These clusters are built for large-scale distributed training of AI models, and they serve a group of about 150 internal researchers whose work covers a diverse set of AI domains, including the full model flow of LLM and VLM training, robotics reinforcement learning (RL) simulation, and post-training for scientific agentic use cases.

Each team had a limit on concurrent GPUs that could be used by workloads which were protected from preemption. Preemptible workloads could exceed that limit on idle GPUs.

Key facts

  • This occurred because researchers โ€” found: they could not launch debugging workloads with low enough latency to tackle problems in real time

Sources & evidence