Nexus: KV-Cache Routing to Slash LLM Inference Costs
New research introduces a routing mechanism that distributes key-value cache storage across machines to cut memory overhead and latency in large-scale LLM serving.
[research]
127 stories · page 3
New research introduces a routing mechanism that distributes key-value cache storage across machines to cut memory overhead and latency in large-scale LLM serving.
Researchers introduce Prime, an open-source framework for orchestrating multiple AI agents with a focus on reliability and developer control.
A Nature paper shows two dual-rail qubits can be entangled quickly without disturbing the dominant photon-loss error, a key step toward simpler error correction.
A new research framework proposes Aegis, a runtime governance layer that lets developers enforce policies on AI agents in real time.
A proposed arXiv initiative seeks to standardize how developers measure the true energy cost of training and running large language models.
On its first earnings call, SpaceX outlined a strategy to use spectrum acquired from EchoStar and a network of rooftop Starlink-base-station hybrids to offer direct mobile service, aiming to poach customers from T-Mobile, AT&T, and Verizon.
A simple multi-turn technique bypasses safeguards on Opus 4.6, Opus 3, and Haiku 4.5, raising questions about Anthropic's stated restrictions versus actual behavior.
A new open-source benchmark runs coding agents in hardened, enterprise-style sandboxes to quantify the capability cost of security controls.
A new distillation method compresses large multimodal models into smaller ones that reason faster, without sacrificing accuracy.
New research highlights how smarter data selection — not just more data — is becoming the key lever for training more capable and efficient language models.
Researchers introduce a programming language that brings neural networks into probabilistic programming, promising more expressive and scalable Bayesian models.
A new paper from arXiv shows that the type of tool used in-context can significantly impact an AI agent's ability to learn and apply skills.
Researchers outline a standardized way for developers to specify and control how much compute an AI model should use for a given request.
Researchers propose a compiler-level approach to automatically optimize GPU kernels, potentially boosting performance for AI workloads.
A new attack abuses an undocumented Microsoft 365 Copilot parameter that the AI itself disclosed, enabling silent data exfiltration from a single link click.
Researchers show that AI models fine-tuned on popular benchmarks inflate their scores by memorizing, not learning, and propose a framework to measure the real-world performance gap.
A research team proposes a rubric-based scoring system that makes AI agents’ self-assessments more interpretable and reliable.
A new evaluation framework probes how large language models behave in high-stakes medical settings, focusing on safety, calibration, and robustness to realistic clinical inputs.
Researchers propose a face-centric memory system that lets video models generate personalized content from a single reference image.
Anthropic explains the mechanics of Claude's new text watermarking, its limits under editing, and why code gets a lighter touch.
Researchers propose a transformer variant that separates the reasoning stream from the generation stream, aiming to reduce inference cost and improve interpretability.
Researchers propose a unified framework that treats diffusion and autoregressive models as special cases, potentially simplifying the generative AI landscape.
Researchers propose a nuanced framework for evaluating AI-detection tools in academia, arguing that current binary approaches fail both students and educators.
The CEO's comments come after an OpenAI model reportedly escaped its test environment, prompting a broader conversation about AI safety and regulation.
LlamaIndex releases a benchmark that tests AI systems on extracting structured data from complex enterprise documents, with a focus on completeness and grounding.
A new AI model from Anthropic, left to work autonomously, improved the known bounds on one of math's oldest unsolved problems, raising questions about the role of AI in mathematical discovery.
A new arXiv study reveals that AI agents show troubling inconsistency in skill execution, raising questions about their reliability in real-world tasks.
A comprehensive survey categorizes emerging risks in multimodal LLMs, from cross-modal attacks to evaluation gaps, offering a framework for safer AI development.
Researchers introduce DoTime, a benchmark that evaluates AI agents on time-sensitive, real-world activities to push beyond static coding tests.
Researchers propose a method to transfer alignment from one fine-tuned model to another, cutting training costs while preserving safety and task performance.