[research]By ByteBulletin Editor
Nexus: KV-Cache Routing to Slash LLM Inference Costs
New research introduces a routing mechanism that distributes key-value cache storage across machines to cut memory overhead and latency in large-scale LLM serving.
[tag]
3 stories
New research introduces a routing mechanism that distributes key-value cache storage across machines to cut memory overhead and latency in large-scale LLM serving.
Researchers propose a transformer variant that separates the reasoning stream from the generation stream, aiming to reduce inference cost and improve interpretability.
New research proposes sticky routing for mixture-of-experts models, keeping tokens on the same expert across layers to cut communication overhead and improve throughput.