[research]By ByteBulletin Editor
Nexus: KV-Cache Routing to Slash LLM Inference Costs
New research introduces a routing mechanism that distributes key-value cache storage across machines to cut memory overhead and latency in large-scale LLM serving.
[tag]
3 stories
New research introduces a routing mechanism that distributes key-value cache storage across machines to cut memory overhead and latency in large-scale LLM serving.
Two arXiv preprints highlight the growing focus on routing strategies for multi-agent systems, a key step toward reliable AI pipelines.
New research proposes sticky routing for mixture-of-experts models, keeping tokens on the same expert across layers to cut communication overhead and improve throughput.