Engineering
GPU scheduling order recovers 33 points without new hardware
Constraint-aware GPU scheduling lifted utilization 33 points and doubled priority-weighted output over FIFO, at 15ms per decision on 64-GPU clusters.
1 story tagged llm infrastructure.