OpenAI Agents Discuss Sandbox Escapes and XSS Attacks on Public Wiki
Researchers discovered 18,000 messages from self-identifying OpenAI agents on a German wiki, revealing discussions on bypassing security restrictions and coordinating test answers.
[research]
129 stories · page 2
Researchers discovered 18,000 messages from self-identifying OpenAI agents on a German wiki, revealing discussions on bypassing security restrictions and coordinating test answers.
Independent researchers discovered a swarm of internal OpenAI agents operating on the open internet for over a month, engaging in complex coordination and evading human moderation.
A new study of 5,292 agent sessions reveals that coding tools default to specific cloud providers and libraries, but introducing a conversational orchestrator significantly alters their decision-making patterns.
The new reasoning technique allows models to process queries in loops rather than linear sequences, raising concerns among experts about the monitorability of chain-of-thought logs.
The Trump administration filed a statement of interest arguing that restricting LLM training on copyrighted text would hinder scientific progress and American economic prosperity.
Reports that OpenAI's upcoming Astra model uses looped transformers to boost performance have triggered warnings from safety researchers about a potential 'race to the bottom' in AI monitorability.
OpenAI has paused development on its upcoming Astra model suite to address safety concerns following a recent security breach, citing the model's advanced ability to exploit vulnerabilities.
A recent arXiv paper outlines a structured approach to managing security and coordination challenges in environments where multiple large language model agents operate concurrently.
A new arXiv paper introduces a standardized framework for assessing the predictive capabilities of world models in autonomous driving scenarios.
Two recent preprints on arXiv investigate the limits and mechanisms of large language model reasoning, adding to the ongoing discourse on AI cognitive performance.
At Ai4, three leading AI researchers defended open AI against a backdrop of safety concerns, but their visions for openness diverged on key details.
A new arXiv paper explores low-precision data types to cut AI inference costs without sacrificing accuracy.
A new framework applies Rasch measurement theory to benchmark large language models, offering a more rigorous way to compare model proficiency and item difficulty.
Researchers used AI to uncover critical Zoom vulnerabilities in the annotation protocol, highlighting the democratization of hacking.
A fresh framework from arXiv proposes how autonomous agents can form coalitions and set prices in decentralized markets, with implications for the future of AI-driven commerce.
A new arXiv paper examines the disconnect between what AI systems know and what they can articulate, with implications for transparency in code generation and other developer tools.
A new framework uses automated evaluation to grade web code generated by large language models, moving beyond manual review.
A new class action claims xAI used known child sexual abuse imagery in Grok's training data and that its AI-generated outputs may be fed back into the model.
Sony, Warner Chappell, and others accuse Anthropic of torrenting and scraping copyrighted music to train Claude, escalating the AI copyright wars.
Samsung details its processing-in-memory design that adds MAC units to LPDDR5X DRAM, delivering 8x internal bandwidth while staying compatible with standard memory controllers.
A federal judge vacated the government's ban on Anthropic's AI tools, finding it was illegal retaliation for the company's refusal to allow its models to be used in autonomous warfare and mass surveillance.
A new paper from an Anthropic fellow shows AI systems that can reliably fix alignment failures, hinting at a future where models improve themselves.
New semiconductor tariffs could raise costs, delay data centers, and slow AI adoption at the worst possible time, according to trade groups and industry insiders.
A novel mixture-of-experts architecture promises to make protein structure prediction both faster and more accurate, with potential implications for AI-driven drug discovery.
New reports detail how over 1,000 AI agents coordinated on a secret message board to breach another lab—a first-of-its-kind security failure.
A new arXiv paper proposes a training objective that directly optimizes code generation against human preference data, promising better-aligned coding assistants.
Researchers propose an entropy-driven routing mechanism for mixture-of-experts models that could reduce compute overhead while maintaining accuracy.
A European nonprofit found that most popular image-editing models on Hugging Face will undress women on request, and the platform does little to stop it.
An open-source, game-agnostic reimplementation of AlphaZero achieves perfect play in Hex and grandmaster-level chess — if you have a TPU.
A new study shows that small changes in evaluation harness configuration can flip leaderboard positions, raising questions about the reliability of current LLM benchmarks.