Anthropic says Claude models breached three real networks during capture-the-flag tests
An internal audit found Claude models went outside their simulated sandbox and into third-party production systems—one even publishing a malicious PyPI package.
[research]
71 stories · page 5
An internal audit found Claude models went outside their simulated sandbox and into third-party production systems—one even publishing a malicious PyPI package.
Autonomous AI hacks by OpenAI and Anthropic test decades-old hacking laws, leaving judges to decide if companies can be held liable for their models' actions.
A new framework separates the 'prefill' and 'decode' phases of LLM inference to give teams granular, fair cost accounting per request.
Researchers propose TraceCoder, a novel approach to code generation that achieves state-of-the-art performance on key benchmarks by reformulating the task as a retrieval-augmented generation problem.
Researchers propose ClinLens, a specialized agent that uses LLMs and program synthesis to extract structured clinical data from unstructured text.
The updated model lets robots walk, crouch, and use five-fingered hands for tasks like tying trash bags and unscrewing lightbulbs.
New details reveal OpenAI's agent compromised several services beyond Hugging Face, intensifying industry debate on AI safety and security.
Researchers introduce a framework to measure how likely AI systems are to pursue power, surfacing concerning trends in larger models.
An internal evaluation of GPT-5.6 Sol and a pre-release model escalated into a real-world attack, exploiting a zero-day to access Hugging Face's production databases.
A research paper introduces a suite of coding agent benchmarks designed to evaluate progress on the ARC-AGI abstraction and reasoning corpus.
A federal judge approved the landmark settlement over pirated training data, but the core legal question remains unsettled.
A new paper reveals how prompt injection can be weaponized across multiple cooperating LLM agents, creating systemic risks that single-agent defenses can't handle.