Latency-Aware Routing Dynamically Balances Speed and Quality Across LLMs
A new routing algorithm uses real-time latency predictions to choose between language models, optimizing for both response time and output quality.
[archive]
292 stories · newest first · page 20
A new routing algorithm uses real-time latency predictions to choose between language models, optimizing for both response time and output quality.
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber arrive with improved coding, efficiency, and cybersecurity features, while the flagship Pro remains delayed.
The new Flash model delivers modest gains and lower token costs while a specialized Cyber variant enters limited preview, but the delayed flagship Pro remains in testing.
Go Micro treats agents as distributed systems, offering a unified runtime for services, agents, and durable workflows with built-in tooling, memory, and cross-framework protocols.
The acquisition, previously undisclosed at this price, signals Netflix's aggressive push into generative AI for content production.
China's AI labs release Kimi K3 and Qwen3.8, touting open-source access and competitive performance at lower cost.
A research paper introduces a suite of coding agent benchmarks designed to evaluate progress on the ARC-AGI abstraction and reasoning corpus.
A federal judge approved the landmark settlement over pirated training data, but the core legal question remains unsettled.
A new paper reveals how prompt injection can be weaponized across multiple cooperating LLM agents, creating systemic risks that single-agent defenses can't handle.
A new server chip, reportedly six to ten times more token-efficient than Google's current hardware, is planned for 2028.
The Model Context Protocol is moving to a stateless session model, reducing infrastructure headaches for companies deploying AI agents at scale.
The startup aims to democratize AI chip access by using an autonomous agent to generate low-level kernel code for any hardware.