Entity

NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference

Disaggregated LLM inference forces the KV cache to traverse the datacenter network before decoding begins, so transfer time enters directly into the Time to First Token (TTFT) budget. Current schedulers route on compute load and prefix-cache locality alone, ignoring the topological distance and dynamic congestion between prefill and decode instances. We close this gap with a thin operator-to-scheduler interface, the network cost oracle, and we prove that ignoring the network term renders cache-a

Paper · arXiv

cs.PF

Authors: Mubarak Adetunji Ojewale
Published: 2026-06-02
Categories: cs.PFcs.AIcs.DCcs.NI

Abstract ↗

via arXiv · 2606.0391