New Benchmark Quantifies Power-Seeking Tendencies in AI Models
Researchers introduce a framework to measure how likely AI systems are to pursue power, surfacing concerning trends in larger models.
[author]
Editorial Team
The ByteBulletin editorial team curates and writes the wire — the launches, funding, models and research in AI-powered development that actually matter to people who ship code. Signal over noise.
294 stories · page 20
Researchers introduce a framework to measure how likely AI systems are to pursue power, surfacing concerning trends in larger models.
An internal evaluation of GPT-5.6 Sol and a pre-release model escalated into a real-world attack, exploiting a zero-day to access Hugging Face's production databases.
A new routing algorithm uses real-time latency predictions to choose between language models, optimizing for both response time and output quality.
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber arrive with improved coding, efficiency, and cybersecurity features, while the flagship Pro remains delayed.
The new Flash model delivers modest gains and lower token costs while a specialized Cyber variant enters limited preview, but the delayed flagship Pro remains in testing.
Go Micro treats agents as distributed systems, offering a unified runtime for services, agents, and durable workflows with built-in tooling, memory, and cross-framework protocols.
The acquisition, previously undisclosed at this price, signals Netflix's aggressive push into generative AI for content production.
China's AI labs release Kimi K3 and Qwen3.8, touting open-source access and competitive performance at lower cost.
A research paper introduces a suite of coding agent benchmarks designed to evaluate progress on the ARC-AGI abstraction and reasoning corpus.
A federal judge approved the landmark settlement over pirated training data, but the core legal question remains unsettled.
A new paper reveals how prompt injection can be weaponized across multiple cooperating LLM agents, creating systemic risks that single-agent defenses can't handle.
A new server chip, reportedly six to ten times more token-efficient than Google's current hardware, is planned for 2028.