From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model
A multimodal training engineer traces low GPU utilization to serial data loading, then shows how concurrency, prefetching, Ray object references, and distributed scheduling shift the bottleneck as training scales.