← Back to feed
AI SecurityEmerging1 sourceSep 9, 2026 · 19:09via AWS Security Blog

The state of AI for security: Measuring what matters most for building trust

Brief

Security teams are starting to actively use AI for security work, including vulnerability triage, penetration testing, threat modeling, incident response, and code review. The promise is speed, but a security tool that moves fast and raises too many false alarms doesn’t save time. Engineers spend time on false alarms, on-call is noisier, and teams distrust findings that matter.

Today, we’re releasing Deception Benchmark, the first benchmark designed to measure that trust problem directly. It tests whether a model can distinguish real vulnerabilities from code that looks risky but is actually safe. The benchmark includes 14,822 samples across 16 languages and more than 70 Common Weakness Enumeration (CWE) categories.

We evaluated 12 models from five providers and are releasing the dataset and whitepaper to the community.

Read more on AWS Security Blog→