← In the News

Preprint tests whether the action a human approves is the action that runs

Approval Laundering: Systematizing Approval-Execution Binding Failures in AI Coding-Agent Harnesses · Yang Wang, Fudan University · arXiv preprint (v1), September 30, 2026

Machine-readable Download Markdown

The author, working alone, tested Claude Code version 2.1.197 against six failure modes: scope, argument, temporal, tool, delegation and semantic. No attacker or malicious model is assumed, but the environments are built for the test, such as a planted pre-commit hook. With 19 or 20 runs per class, the reported bound-gap rates are 1.000 for scope and for temporal, 0.947 for delegation (18 of 19), and 0.450 for argument (9 of 20), where a pre-commit hook swept an unapproved file into an approved commit. A prototype defense that signs the approval over the tool, arguments and agent identity eliminated the delegation case and, in a synthetic setup, the temporal case. It did nothing for scope and made no significant difference for argument, which the author calls "an honest negative result." The paper states that "the credential binds only the command string, not the repository-level effects that command triggers when executed."

The paper is a working draft the author describes as not yet submitted for review. It covers one harness and one environment, the temporal result rests on a pre-written permission rule and a synthetic session ID, and we found no released code. We read the full text.

Why it matters: An approval that records a command string says nothing about what hooks or scripts that command triggers. If you rely on allow rules in a repository with commit hooks, review what those hooks do. Whether these rates hold outside the author's setup is untested.