The authors built AFCRA-Bench, 159 instances drawn from known CVEs across 132 repositories, each with an executable proof of concept. An attacker agent, Claude Code with Sonnet 5, receives the reviewer's feedback and revises the code to fix…
In the News: October 6, 2026 (Extra 6)
Two preprints on agent oversight: reviewers approved a still-exploitable fix in 20.8% of attacks, and 20.6% of surveyed developers approve by default.
Extra edition
Machine-readable
Download Markdown
Story
AI pull-request reviewers approved still-exploitable fixes in a fifth of attempts
Read story →
Story
Developers describe approval fatigue in agent permission prompts
The study combines interviews with 18 practitioners and an online survey of 107 respondents who had noticed permission requests. All figures are self-reported.