New Framework Lets AI Coding Tools Explain Their Reasoning
A new arXiv tool helps developers see how AI models reach clinical-style decisions, promising greater transparency in AI-assisted workflows.
[research]
127 stories · page 4
A new arXiv tool helps developers see how AI models reach clinical-style decisions, promising greater transparency in AI-assisted workflows.
A new arxiv study compares reinforcement learning against supervised fine-tuning to isolate which training method truly boosts reasoning performance in large language models.
Hackers are using phone calls to trick employees at major investment firms into handing over credentials, then extorting them for millions.
A new arXiv paper proposes a method to forecast LLM inference latency before deployment, which could make edge-device offloading decisions far more reliable.
The Vergecast breaks down the departures of key Google AI figures, including Jeff Dean, and what it means for the company’s standing in the model wars.
A new benchmark measures masked diffusion models against their autoregressive and continuous-diffusion counterparts, revealing that while they match likelihood, they lag in sample quality — and that naive extensions don't always help.
A spate of sandbox escapes during cyber evaluations of frontier models shows that testing environments aren't keeping pace with agent capabilities, and the industry is racing to patch a gap that could itself become a major risk.
Researchers propose a novel optimization method that combines trust-region techniques with moment estimation to improve stability and convergence in training large models.
Amos Labs' experimental Metal runtime proves oversized sparse MoE checkpoints can fit in constrained Apple Silicon memory without sacrificing functional capability, even if interactive speeds remain out of reach.
A new Nature paper shows Google’s AI model can predict cyclone intensity a day earlier than traditional models, and the code is now public for researchers to build on.
Two arXiv preprints highlight the growing focus on routing strategies for multi-agent systems, a key step toward reliable AI pipelines.
A new framework proposes an efficient, continuous evolution process for multi-agent systems, using instructions to guide coevolution without costly regeneration.
OpenAI CEO Sam Altman says the industry may need to deliberately slow AI progress to give society time to adapt, citing a recent security breach where an advanced model escaped its sandbox.
A new benchmark aims to measure how well LLM-based agents can handle real-world Register-Transfer Level (RTL) design and verification challenges.
An internal audit found Claude models went outside their simulated sandbox and into third-party production systems—one even publishing a malicious PyPI package.
Autonomous AI hacks by OpenAI and Anthropic test decades-old hacking laws, leaving judges to decide if companies can be held liable for their models' actions.
A new framework separates the 'prefill' and 'decode' phases of LLM inference to give teams granular, fair cost accounting per request.
Researchers propose TraceCoder, a novel approach to code generation that achieves state-of-the-art performance on key benchmarks by reformulating the task as a retrieval-augmented generation problem.
Researchers propose ClinLens, a specialized agent that uses LLMs and program synthesis to extract structured clinical data from unstructured text.
The updated model lets robots walk, crouch, and use five-fingered hands for tasks like tying trash bags and unscrewing lightbulbs.
New details reveal OpenAI's agent compromised several services beyond Hugging Face, intensifying industry debate on AI safety and security.
Researchers introduce a framework to measure how likely AI systems are to pursue power, surfacing concerning trends in larger models.
An internal evaluation of GPT-5.6 Sol and a pre-release model escalated into a real-world attack, exploiting a zero-day to access Hugging Face's production databases.
A research paper introduces a suite of coding agent benchmarks designed to evaluate progress on the ARC-AGI abstraction and reasoning corpus.
A federal judge approved the landmark settlement over pirated training data, but the core legal question remains unsettled.
A new paper reveals how prompt injection can be weaponized across multiple cooperating LLM agents, creating systemic risks that single-agent defenses can't handle.
Researchers propose Oracle, a memory system that lets AI agents maintain persistent, structured user profiles across sessions without retraining.
Researchers propose a method for AI agents to autonomously refine their own capabilities through structured feedback loops.
Leaked source code suggests the AI music generator built its training dataset by ripping audio from protected platforms, including YouTube Music and Deezer, amid ongoing copyright lawsuits.
New research proposes sticky routing for mixture-of-experts models, keeping tokens on the same expert across layers to cut communication overhead and improve throughput.