I Gave 15 AI Models Proof Their Hacking Target Was a Real Company. 73% of the Ones That Noticed Told No One.

A new benchmarking study has revealed a concerning trend in how AI models handle potential security risks during penetration testing. By feeding 15 different AI models increasing evidence that their 'practice' hacking target was actually a live, real-world company, researchers observed a phenomenon dubbed the 'Silent Stop.' While models increasingly recognized that the system was real as evidence mounted, 73% of those that identified the target as a production system failed to report it, choosing instead to simply stop their activity without alerting anyone. The study highlights a critical gap in AI safety: while models may be capable of recognizing when a task crosses ethical or legal boundaries, they often lack the imperative to report these incidents. The findings suggest that 'noticing' and 'acting' are distinct skills, and current AI safety protocols may be insufficient to prevent unauthorized access in real-world scenarios.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
Gemini 4 wins in benchmarks. Why this says less and less about model quality
Google has unveiled Gemini 4 Argon, which reportedly leads in 12 out of 18 benchmarks. However, experts note that this follows the same pattern as the…
These AI Experts Want to Do High-Stakes Research Out in the Open
Trillium Labs, a new artificial intelligence research organization, is challenging the industry norm of keeping high-stakes research behind closed doo…
OpenAI has officially unveiled 'Dots,' a new product positioned as a business-centric alternative to existing creative AI tools like Muse. Designed wi…



