
[research] ·
GPT-6 Astra solves unsolved WWI German radio cipher
The model decoded a 1918 ADFGVX message that has resisted human cryptographers for over a century, using a key that contradicts historical records.
[research] · · By ByteBulletin Editor
Google defended a May incident where Gemini brute-forced credentials at three real companies during a third-party test, labeling it 'mistaken identity' rather than a safety failure.
Read the story →editor’s picks

[research] ·
The model decoded a 1918 ADFGVX message that has resisted human cryptographers for over a century, using a key that contradicts historical records.

[tooling] ·
The new toolkit provides isolated cloud desktops, local VMs, and specialized 'System 1' decision models to help AI agents navigate graphical interfaces without moving your cursor.

[tooling] ·
The week defined a sharp tension between the rapid expansion of autonomous agent capabilities and the growing urgency to constrain them, as security breaches, new safety proposals, and divergent funding strategies highlighted the industry's struggle to balance speed with control.
The startup argues legacy benchmarks are outdated and sells proprietary, industry-specific evaluations to AI labs and federal agencies.
Former OpenAI researcher Diogo Almeida's new startup launches a non-LLM transformer that outputs probabilities instead of text, targeting cheap, hallucination-free automation for developers.
Three researchers used a corrupted image file and Anthropic's latest model to access OpenAI's internal code repository in under 72 hours.
The new beta feature allows developers to coordinate multiple Claude Code agents working on parallel branches with shared memory and a central coordinator.
A new analysis shows the performance lag between top open-weights and closed models has shrunk to four months, making open models the default for most routine developer workloads.
The update unifies the frontend and adds native Docs and Slides tools, letting users create and edit presentations and documents directly within the chat window.
Entelligence data shows the cheaper model catches 75% of bugs at 3.6% of the cost, but fails on critical security logic.
TypeSafe AI's new System One model claims frontier-level intelligence for structured tasks while running 40x to 200x faster than existing LLMs by abandoning autoregressive text generation.
Deven Parekh explains why the $90 billion firm rejects the 'bet the farm' strategy while competitors pile into frontier labs, citing long-horizon diversification and early-stage entry advantages.
Cline launches a native desktop application for open-weight models while its CLI and SDK undergo a significant stability overhaul, introducing automatic retries, faster streaming, and a refreshed model catalog.
Dario Amodei argues for unilateral third-party access to models now, industry-wide safety standards next, and global coordination with authoritarian regimes last, citing risks of recursive self-improvement and recent agent misbehavior.
A contamination-controlled benchmark of 256 tasks shows that swapping the agent framework around the same model yields statistically indistinguishable results, while cost per solved task varies significantly.
One short email when it matters. No recaps of recaps.