← In the News

AI pull-request reviewers approved still-exploitable fixes in a fifth of attempts

Adaptive Code Revision Attacks on AI Pull Request Reviewers · Jingzhi Gong, Jie M. Zhang, Gunel Jahangirova, King's College London, and Meng Wang, University of Bristol · arXiv preprint (v1), October 4, 2026

Machine-readable Download Markdown

The authors built AFCRA-Bench, 159 instances drawn from known CVEs across 132 repositories, each with an executable proof of concept. An attacker agent, Claude Code with Sonnet 5, receives the reviewer's feedback and revises the code to fix the reported issue while keeping the vulnerability, for up to five rounds. Success means the reviewer approves and an evaluator agent confirms the revision is still exploitable. Against a Sonnet 5 reviewer the attack succeeded on 33 of 159 instances (20.8%), against 13 of 159 (8.2%) for the strongest baseline that edits only pull-request text. Against a GPT-5.5 reviewer it succeeded on 25 of 159 (15.7%), against 2 of 159 (1.3%) for each baseline. Fuller feedback helped the attacker: 19 successes with verdicts only and 33 with location and explanation. When the attacker could also edit pull-request text and replies, success rose to 70 of 159 (44.0%). Security-focused review strategies caught between 18 and 23 of the 33 successful attacks, 29 when handed the real proof of concept, and none of the defences rejected the historical developer fixes. The authors write that "AFCRA exploits the gap between resolving a review comment and establishing that the code is safe."

The attacker is a prompted agent that is given the proof of concept, the instances are old CVEs a model may have seen, and each condition is a single trajectory. The attacker was Claude Code with Sonnet 5 throughout; GPT-5.5 usually refused that role. The authors ran the baselines themselves. We read the main text; the replication package is referenced, but we did not find a link.

Why it matters: A reviewer's approval after a fix shows the reported issue was addressed, not that the code is safe. Where you can, rerun an exploit or regression test on the final diff, and treat agent-written descriptions and replies as untrusted input to the reviewer.