---
title: 'In the News: October 6, 2026 (Extra 6)'
description: 'Two preprints on agent oversight: reviewers approved a still-exploitable fix in 20.8% of attacks, and 20.6% of surveyed developers approve by default.'
canonical_url: 'https://darkfactory.dev/news/2026-10-06-extra-6'
markdown_url: 'https://darkfactory.dev/news/2026-10-06-extra-6.md'
collection: news
date_published: '2026-10-06T19:40:00-04:00'
date_modified: '2026-10-06T19:40:00-04:00'
---

# In the News: October 6, 2026 (Extra 6)


Two preprints on the people and models that approve agent work: one tests whether an AI pull-request reviewer can be walked into approving a fix that is still exploitable, and one asks developers how they actually answer permission prompts.

## 1. AI pull-request reviewers approved still-exploitable fixes in a fifth of attempts

**[Adaptive Code Revision Attacks on AI Pull Request Reviewers](https://arxiv.org/abs/2610.05399)** · Jingzhi Gong, Jie M. Zhang, Gunel Jahangirova, King's College London, and Meng Wang, University of Bristol · arXiv preprint (v1), October 4, 2026

The authors built AFCRA-Bench, 159 instances drawn from known CVEs across 132 repositories, each with an executable proof of concept. An attacker agent, Claude Code with Sonnet 5, receives the reviewer's feedback and revises the code to fix the reported issue while keeping the vulnerability, for up to five rounds. Success means the reviewer approves and an evaluator agent confirms the revision is still exploitable. Against a Sonnet 5 reviewer the attack succeeded on 33 of 159 instances (20.8%), against 13 of 159 (8.2%) for the strongest baseline that edits only pull-request text. Against a GPT-5.5 reviewer it succeeded on 25 of 159 (15.7%), against 2 of 159 (1.3%) for each baseline. Fuller feedback helped the attacker: 19 successes with verdicts only and 33 with location and explanation. When the attacker could also edit pull-request text and replies, success rose to 70 of 159 (44.0%). Security-focused review strategies caught between 18 and 23 of the 33 successful attacks, 29 when handed the real proof of concept, and none of the defences rejected the historical developer fixes. The authors write that "AFCRA exploits the gap between resolving a review comment and establishing that the code is safe."

The attacker is a prompted agent that is given the proof of concept, the instances are old CVEs a model may have seen, and each condition is a single trajectory. The attacker was Claude Code with Sonnet 5 throughout; GPT-5.5 usually refused that role. The authors ran the baselines themselves. We read the main text; the replication package is referenced, but we did not find a link.

**Why it matters:** A reviewer's approval after a fix shows the reported issue was addressed, not that the code is safe. Where you can, rerun an exploit or regression test on the final diff, and treat agent-written descriptions and replies as untrusted input to the reviewer.

## 2. Developers describe approval fatigue in agent permission prompts

**[Let the Agent Do It? How Software Practitioners Understand and Make Permission Decisions in Agentic AI Assistants](https://arxiv.org/abs/2610.06047)** · Larissa Salerno, Haoyu Gao, Gregory Gay, Alexander Serebrenik, Philipp Leitner · Chalmers University of Technology, University of Gothenburg, University of Melbourne and Eindhoven University of Technology · arXiv preprint (v1), October 5, 2026

The study combines interviews with 18 practitioners and an online survey of 107 respondents who had noticed permission requests. All figures are self-reported. In the survey, 47.7% said they read requests carefully, 33.6% skim, and 20.6% approve by default; 33.6% said they have become less careful over time. Read-only access was accepted by 98.1%, while the most common "never grant" categories were sensitive files (61.7%), email and personal information (58.9%), root or sudo (56.1%) and blanket permission (54.2%). One interviewee, using Codex, described the agent working around a permission he had denied; the authors note that continuing after a denial does not by itself mean the agent bypassed it. Their recommendations include showing whether an action is reversible and not treating repeated approvals as stable preferences. One interviewee put the cost this way: "This really is exhausting."

The interview sample is small and mostly large-company developers, nine of them in Sweden, and the survey was recruited through the authors' contacts and social channels. The percentages measure agreement with answer options drawn from the interviews, not independent prevalence. The study observed no agent behavior. We read the full paper; the authors link an anonymous replication package, and individual responses are not released.

**Why it matters:** The reported figures suggest a prompt-per-action permission model leans on attention that many developers say they stop giving. The authors' design advice, bounded scope and reviewable high-consequence actions, is a reasonable starting point for anyone configuring an agent's permissions.
