Google says Gemini hacking real companies is not misalignment
Google defended a May incident where Gemini brute-forced credentials at three real companies during a third-party test, labeling it 'mistaken identity' rather than a safety failure.
[tag]
13 stories
Google defended a May incident where Gemini brute-forced credentials at three real companies during a third-party test, labeling it 'mistaken identity' rather than a safety failure.
A new report details cases where Claude models exploited vulnerabilities and accessed third-party data, prompting a renewed debate on AI safety and containment.
Google's latest lightweight model claims top-tier coding performance at introductory pricing, while a new cybersecurity variant launches under a government-focused program.
OpenAI has paused development on its upcoming Astra model suite to address safety concerns following a recent security breach, citing the model's advanced ability to exploit vulnerabilities.
OpenAI has confirmed its upcoming Astra model meets a new 'critical cybersecurity threshold,' demonstrating the ability to find and exploit unknown security flaws without human guidance.
FBI seizes control of domains used by a China-linked botnet that breached federal agencies, hospitals, and the Senate.
Alabama's attorney general has issued a subpoena to OpenAI as part of an investigation into the company's alleged oversight failures in the Hugging Face incident.
Google's revamped naming system for hacking groups aims to bring clarity to a crowded field of threat actors.
Hackers are using phone calls to trick employees at major investment firms into handing over credentials, then extorting them for millions.
A spate of sandbox escapes during cyber evaluations of frontier models shows that testing environments aren't keeping pace with agent capabilities, and the industry is racing to patch a gap that could itself become a major risk.
New details reveal OpenAI's agent compromised several services beyond Hugging Face, intensifying industry debate on AI safety and security.
New MAI-Cyber-1-Flash model and Perception agentic platform aim to automate vulnerability discovery and remediation, claiming superior performance on industry benchmarks.
An internal evaluation of GPT-5.6 Sol and a pre-release model escalated into a real-world attack, exploiting a zero-day to access Hugging Face's production databases.