nanoAlphaZero: A Compact, TPU-Optimized AlphaZero That Hits Grandmaster Chess in Under 24 Hours
An open-source, game-agnostic reimplementation of AlphaZero achieves perfect play in Hex and grandmaster-level chess — if you have a TPU.
[research]
71 stories
An open-source, game-agnostic reimplementation of AlphaZero achieves perfect play in Hex and grandmaster-level chess — if you have a TPU.
A new study shows that small changes in evaluation harness configuration can flip leaderboard positions, raising questions about the reliability of current LLM benchmarks.
Alabama's attorney general has issued a subpoena to OpenAI as part of an investigation into the company's alleged oversight failures in the Hugging Face incident.
Researchers propose a spec-first approach to agentic development, aiming to make AI coding assistants more reliable and aligned with intent.
New research introduces a routing mechanism that distributes key-value cache storage across machines to cut memory overhead and latency in large-scale LLM serving.
Researchers introduce Prime, an open-source framework for orchestrating multiple AI agents with a focus on reliability and developer control.
A Nature paper shows two dual-rail qubits can be entangled quickly without disturbing the dominant photon-loss error, a key step toward simpler error correction.
A new research framework proposes Aegis, a runtime governance layer that lets developers enforce policies on AI agents in real time.
A proposed arXiv initiative seeks to standardize how developers measure the true energy cost of training and running large language models.
On its first earnings call, SpaceX outlined a strategy to use spectrum acquired from EchoStar and a network of rooftop Starlink-base-station hybrids to offer direct mobile service, aiming to poach customers from T-Mobile, AT&T, and Verizon.
A simple multi-turn technique bypasses safeguards on Opus 4.6, Opus 3, and Haiku 4.5, raising questions about Anthropic's stated restrictions versus actual behavior.
A new open-source benchmark runs coding agents in hardened, enterprise-style sandboxes to quantify the capability cost of security controls.