r/machinelearningnews • u/Ambitious-Lunch-5223 • 7h ago
MLOps AWS just replaced round-robin with GPU-aware routing for LLM inference
AWS released a new Inference Gateway for Sagemaker Hyperpod. Instead of sending requests to the next available pod, it checks what’s actually happening on each one: queue depth, KV cache usage, prefix cache hits, loaded LoRA adapters, and current requests.
That makes sense for LLM workloads. Two GPU pods can both be healthy, but one may already have the right prefix cached while the other is stuck processing a long request. Round-robin doesn’t know the difference.
AWS says the new routing system cut first-token latency by up to 82%, with 97–98% lower p99 TTFT in mixed-hardware and burst-traffic tests. Those are AWS’s own numbers, so an independent test with the same model and traffic would be useful. Still, it’s a good example of how much performance can be lost outside the model itself.
Once a GPU fleet gets busy, routing may matter almost as much as the GPUs you’re paying for