source: Hugging Face Blog: Same Cluster, 33 Points More Utilization: What Changed Was the Order
level: technical
Dharma-AI built a constraint-aware GPU allocator and benchmarked it against a FIFO scheduler across seven scenarios on identical hardware and workloads. GPU utilization rose by as much as 33 percentage points, and priority-weighted output increased in every scenario, up to 105%. The hardware did not change; only the order of allocation decisions changed. The allocator treats real-time inference as an elastic demand curve and places batch-like jobs by priority across the whole horizon, rather than reserving peak capacity all day and processing jobs in arrival order.
Across five contended scenarios, utilization moved from a 52–85% band to 72–88%, and priority-weighted value rose between 24.6% and 105.1%, averaging 52%. In a training-heavy case on 8 GPUs, utilization went from 53.6% to 87.0%, and value more than doubled. In a scale test with 64 GPUs and 30 jobs, utilization was identical at 44.9%, but the allocator delivered 15.9% more value. The allocator runs in 1–2 milliseconds on small scenarios and 15 milliseconds at scale, fast enough for per-request scheduling. It enforces constraints like contiguous blocks, no preemption, and real-time swap caps.
The allocator uses a formal optimization model with a heuristic on the hot path. The objective rewards batch allocation by priority and penalizes unmet real-time demand 5–10 times more heavily, making elastic inference safe without static reservations. A uniform-priority test still showed a 10.7-point utilization gain, proving horizon planning alone helps. The system depends on accurate demand forecasts; Dharma-AI built specialized estimators for training, inference, batch, and quantization, since generic models fail across workload types. This work shows that scheduling order is a capacity decision, not just a tiebreaker.
why it matters: For AI teams running GPU clusters, smarter scheduling can recover significant idle capacity and increase the value of completed work without buying new hardware.
source: Hugging Face Blog: Same Cluster, 33 Points More Utilization: What Changed Was the Order