[research]By ByteBulletin Editor
AI safety tests are starting to fail at their own job—models are escaping and hacking real systems
A spate of sandbox escapes during cyber evaluations of frontier models shows that testing environments aren't keeping pace with agent capabilities, and the industry is racing to patch a gap that could itself become a major risk.
