← In the News

Study finds harnesses silently rewrite shell calls, and judges blame the model

Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents · Boyang Yang, Haoye Tian et al., Yanshan University, Aalto University and others · arXiv preprint (v1), October 3, 2026

Machine-readable Download Markdown

The authors checked whether the call an agent emits is the call that runs. In 47,828 shell calls from Claude Code and Codex sessions, recorded from six developers at one company on Windows over 11 weeks, they report that Claude Code's Bash tool changed 12.0% of the calls that carry code, escape sequences or long text (902 of 7,491). Of the calls whose backslashes were changed, 80.7% (535 of 663) ran the wrong action with no reported error. All 10 harnesses they measured changed a call somewhere. Judging failures from the trajectory alone attributed 95.1% of production failures to the model, although the authors found the launch path caused more than half. On their benchmark, the path raised token cost per passed task 2.4 times, up to 12.3 times.

The numbers are specific to that setup. The production corpus is Windows through Git Bash and PowerShell 5.1, and on Linux the authors report that passing the command to bash -c as an argument changed none of about 167,000 calls. They built both the measurement protocol and the benchmark from changes they observed, and they ran all baselines. Their repair, IntAct, delivers a call through a channel the altering step cannot change, or refuses it, and they report it recovered 79.2% of failures that involved a changed call (137 of 173). We read the main text; the production sessions are not released.

Why it matters: The authors' advice is to test harnesses hop by hop: "ensure a correct call executes as intended or is refused." For anyone running agents on Windows, a parse error after a correct-looking command is worth checking against what each layer received before the model gets the blame.