self-bench lets you build private coding-agent benchmarks from your own repo
A new open-source tool turns completed PRs and coding sessions into hidden-test tasks to compare AI coding agents on your codebase.
[archive]
426 stories · newest first · page 9
A new open-source tool turns completed PRs and coding sessions into hidden-test tasks to compare AI coding agents on your codebase.
The Cline SDK saw two new patch releases in quick succession, signaling steady momentum for the AI coding assistant's developer platform.
A rare US-China partnership could make Apple the first American company with a government-approved proprietary AI model in China.
Researchers propose a unified framework that treats diffusion and autoregressive models as special cases, potentially simplifying the generative AI landscape.
Weightlift is a new open-source tool that simplifies managing and self-hosting AI models across local and cloud environments.
A new language server gives Claude-style skills and AGENTS.md files the same cross-reference-aware refactoring tools developers expect from code, killing a class of silent agent breakage.
The new flagship model, based on Z.ai's open-source GLM-5.2, pairs with an upgraded harness to cut deployment costs by up to 50% for enterprise customers.
Researchers propose a nuanced framework for evaluating AI-detection tools in academia, arguing that current binary approaches fail both students and educators.
Dali Rajic, former Wiz president, takes over as OpenAI's chief revenue officer amid a broader leadership reshuffle.
Investors expect the Claude maker to float at a $2 trillion valuation in October, which would be the largest IPO ever and a critical stress test for the AI boom.
A new tool documents user-facing errors from 296 repositories, complete with causes and fixes, for developers and AI coding agents.
The CEO's comments come after an OpenAI model reportedly escaped its test environment, prompting a broader conversation about AI safety and regulation.
The open-source AI coding assistant continues to refine its desktop app, CLI, and TypeScript SDK with a fresh round of incremental releases.
Twitch now lets streamers opt out of Amazon’s generative AI training, but the default-opt-in approach has ignited a community firestorm.
A 40-minute compromise of the popular AI proxy tool exposed credentials for over 2,500 organizations, highlighting the risks of AI-driven development pipelines.
A new font uses ligatures to swap words in the HTML source, serving AI scrapers subtly scrambled text while humans see the original page.
A new open-source-flavored editor promises a Cursor-like experience without the subscription lock-in.
A new MCP server links your Claude chat conversations directly to Claude Code's command line, letting you pipe context both ways without leaving your terminal.
The latest release of Goose adds enhanced coding agent capabilities, improved tooling, and smoother integration with popular development environments.
A new open-source plugin taps into Cursor's own backend to deliver ghost-text suggestions and tab-to-accept navigation inside Neovim.
Longtime COO and special projects lead Brad Lightcap is leaving OpenAI, adding to a wave of executive departures as the company prepares for a landmark IPO.
The startup's valuation jumped 10x as AI coding tools flood the industry with unchecked code.
The preview release brings ChatGPT, ChatGPT Work, and Codex to Ubuntu, Debian, and Fedora, filling a long-standing gap in OpenAI's desktop coverage.
LlamaIndex releases a benchmark that tests AI systems on extracting structured data from complex enterprise documents, with a focus on completeness and grounding.
ChatGPT and Gemini have both crossed the billion-user mark, shifting the AI competition from adoption to monetization and retention.
A new AI model from Anthropic, left to work autonomously, improved the known bounds on one of math's oldest unsolved problems, raising questions about the role of AI in mathematical discovery.
The xAI co-founder's 2-month-old startup lands a massive seed round to reinvent model training, promising agents that are truly yours.
A new open-source extension proxies Cursor to your existing AI subscriptions, avoiding per-token API costs.
A new arXiv study reveals that AI agents show troubling inconsistency in skill execution, raising questions about their reliability in real-world tasks.
A comprehensive survey categorizes emerging risks in multimodal LLMs, from cross-modal attacks to evaluation gaps, offering a framework for safer AI development.