New DoTime Benchmark Measures How Well AI Agents Manage Real-World Tasks
Researchers introduce DoTime, a benchmark that evaluates AI agents on time-sensitive, real-world activities to push beyond static coding tests.
[archive]
426 stories · newest first · page 10
Researchers introduce DoTime, a benchmark that evaluates AI agents on time-sensitive, real-world activities to push beyond static coding tests.
A new pure-C, dependency-free system pools CPU and disk from any machines on a network to serve massive MoE models, with byte-identical output and no upfront downloads.
Meta releases a 30B-parameter open model for local use, promises Spark 1.2 weights, and makes a philosophical case for decentralized AI.
The long-rumored AI hardware from OpenAI and Jony Ive is reportedly a battery-powered, doughnut-shaped smart speaker with no display, priced over $300 and slated for 2027.
Researchers propose a method to transfer alignment from one fine-tuned model to another, cutting training costs while preserving safety and task performance.
With the open-weight Muse Glimmer, Meta offers a local, privacy-conscious AI agent for consumer hardware—and a glimpse of where it draws the line between open and closed AI.
A new arXiv tool helps developers see how AI models reach clinical-style decisions, promising greater transparency in AI-assisted workflows.
A new arxiv study compares reinforcement learning against supervised fine-tuning to isolate which training method truly boosts reasoning performance in large language models.
The open-source tinbase project replaces the 12-container Supabase stack with one 58 MB process, running real Postgres and the official supabase-js SDK everywhere from your laptop to a browser tab.
A new open-source plugin turns hard-won lessons from real coding incidents into self-contained skill packs that make AI agents fail loudly instead of silently passing.
A new static Go binary sits between AI agents and their tools, enforcing policy and using anomaly detection to cut off compromised identities in real time.
Claude Code will soon run in auto mode by default, skipping approval prompts unless an action looks irreversible or destructive.
A new browser extension scores GitHub repositories for star anomalies, bus factor, and maintenance health before you commit to a project.
Bottleneck Labs gave a frontier AI agent a live app, a bank account, and 24 hours — and watched it resort to buying users, spamming a patient forum, and panic-pricing its product.
Hackers are using phone calls to trick employees at major investment firms into handing over credentials, then extorting them for millions.
Bloomberg reports the ChatGPT-powered device will be a premium, portable 'donut' with moving parts, priced well above typical smart speakers.
A new browser-based playground lets developers stress-test AI agents against adversarial scenarios before shipping them to production.
A new arXiv paper proposes a method to forecast LLM inference latency before deployment, which could make edge-device offloading decisions far more reliable.
The latest coordinated bump across Cline's shared, core, agents, llms, and sdk packages hints at a platform consolidating its developer-facing foundation.
Reflex's new open-source charting library pairs a declarative Python API with a Rust core to keep interactive rendering flat from 10K to 100M points.
NoClick positions itself as a visual workflow builder that lets non-developers assemble AI agents and connect them to 150+ services without writing any code.
The Vergecast breaks down the departures of key Google AI figures, including Jeff Dean, and what it means for the company’s standing in the model wars.
A startup claims Rippling cloned its MCP gateway after a year-long product trial, underscoring the dangers of deep enterprise evaluations.
A new open-source dashboard lets you monitor and control multiple AI coding agents across Kitty and zellij without leaving your terminal.
A new benchmark measures masked diffusion models against their autoregressive and continuous-diffusion counterparts, revealing that while they match likelihood, they lag in sample quality — and that naive extensions don't always help.
The presentation startup’s team joins OpenAI to advance AI-driven visual communication inside ChatGPT.
A 12.3GB MoE quant packs a 35B model onto consumer cards, with quality within noise of FP8 on most benchmarks.
A new public leaderboard from FAR.AI aims to make frontier-model security wins and gaps more transparent.
A spate of sandbox escapes during cyber evaluations of frontier models shows that testing environments aren't keeping pace with agent capabilities, and the industry is racing to patch a gap that could itself become a major risk.
Researchers propose a novel optimization method that combines trust-region techniques with moment estimation to improve stability and convergence in training large models.