AI struggles to patch vulns without adult supervision
What happened
Researchers tested AI models that generate security patches and found autonomous fixes fully remediated vulnerabilities in only a minority of cases. The study produced thousands of candidate patches and concluded many fixes either failed, changed application behaviour, or introduced new issues. Procurement should treat AI autopatching as an assistive output that requires gating, test evidence, and human sign‑off before deployment
Why the category manager should care
Treat AI patch outputs as preliminary artifacts; require human review, acceptance tests, and rollback guarantees before a vendor or tool can affect production
Key facts
- Study produced several thousand candidate patches from two frontier models
- Only about a quarter of AI patches fully resolved vulnerabilities without altering app behaviour
- Large share of patches were fragile or introduced new security issues