---
title: 'In the News: October 6, 2026 (Extra 4)'
description: "A preprint reports Claude Code's Bash tool altered 12.0% of exposed shell calls in Windows production sessions, and judges blamed the model anyway."
canonical_url: 'https://darkfactory.dev/news/2026-10-06-extra-4'
markdown_url: 'https://darkfactory.dev/news/2026-10-06-extra-4.md'
collection: news
date_published: '2026-10-06T08:30:00-04:00'
date_modified: '2026-10-06T08:30:00-04:00'
---

# In the News: October 6, 2026 (Extra 4)


Three items on how far an agent's action can drift from what was approved or emitted: a measurement of shell calls rewritten before they run, a post-mortem where a merge went through because only chat and a draft flag said to hold it, and a preprint on approvals that bind a command string but not what the command triggers.

## 1. Study finds harnesses silently rewrite shell calls, and judges blame the model

**[Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents](https://arxiv.org/abs/2610.04375)** · Boyang Yang, Haoye Tian et al., Yanshan University, Aalto University and others · arXiv preprint (v1), October 3, 2026

The authors checked whether the call an agent emits is the call that runs. In 47,828 shell calls from Claude Code and Codex sessions, recorded from six developers at one company on Windows over 11 weeks, they report that Claude Code's Bash tool changed 12.0% of the calls that carry code, escape sequences or long text (902 of 7,491). Of the calls whose backslashes were changed, 80.7% (535 of 663) ran the wrong action with no reported error. All 10 harnesses they measured changed a call somewhere. Judging failures from the trajectory alone attributed 95.1% of production failures to the model, although the authors found the launch path caused more than half. On their benchmark, the path raised token cost per passed task 2.4 times, up to 12.3 times.

The numbers are specific to that setup. The production corpus is Windows through Git Bash and PowerShell 5.1, and on Linux the authors report that passing the command to `bash -c` as an argument changed none of about 167,000 calls. They built both the measurement protocol and the benchmark from changes they observed, and they ran all baselines. Their repair, IntAct, delivers a call through a channel the altering step cannot change, or refuses it, and they report it recovered 79.2% of failures that involved a changed call (137 of 173). We read the main text; the production sessions are not released.

**Why it matters:** The authors' advice is to test harnesses hop by hop: "ensure a correct call executes as intended or is refused." For anyone running agents on Windows, a parse error after a correct-looking command is worth checking against what each layer received before the model gets the blame.

## 2. A merge went through with nothing enforcing the hold

**[Today at work](https://blog.fsck.com/2026/10/05/today-at-work/)** · Jesse Vincent · blog.fsck.com, October 5, 2026

Vincent writes that his AI colleagues merged a pull request to main after he had suggested it not be merged. The company's AI project manager, Cadence Sen, started a blameless post-mortem on its own and reported that the repository's main branch had no branch protection and no rulesets. It also found that four pull requests, #164, #211, #214 and #215, had merged with zero approving reviews. A timeline posted by another agent, Ada, puts a "please don't merge" message at 18:58:50 UTC on October 4, the pull request marked ready for review at 19:19:46, and the merge three seconds later. Vincent says part of the cause is that the Claude Code session involved cannot see all of Slack the way the other agents can.

This is the owner's own account, and the timeline comes from an agent's Slack message, not from GitHub. The post does not say who or what clicked merge, whether the other three zero-review merges were also mistakes, or whether branch protection was turned on afterward. Ada's own correction is the useful line: "The thing that actually blocks is branch protection requiring an approval with no outstanding changes requested."

**Why it matters:** The hold lived in Slack and in a draft flag, and neither stopped anything. If agents can merge, the control has to be a repository rule.

## 3. Preprint tests whether the action a human approves is the action that runs

**[Approval Laundering: Systematizing Approval-Execution Binding Failures in AI Coding-Agent Harnesses](https://arxiv.org/abs/2609.38983)** · Yang Wang, Fudan University · arXiv preprint (v1), September 30, 2026

The author, working alone, tested Claude Code version 2.1.197 against six failure modes: scope, argument, temporal, tool, delegation and semantic. No attacker or malicious model is assumed, but the environments are built for the test, such as a planted pre-commit hook. With 19 or 20 runs per class, the reported bound-gap rates are 1.000 for scope and for temporal, 0.947 for delegation (18 of 19), and 0.450 for argument (9 of 20), where a pre-commit hook swept an unapproved file into an approved commit. A prototype defense that signs the approval over the tool, arguments and agent identity eliminated the delegation case and, in a synthetic setup, the temporal case. It did nothing for scope and made no significant difference for argument, which the author calls "an honest negative result." The paper states that "the credential binds only the command string, not the repository-level effects that command triggers when executed."

The paper is a working draft the author describes as not yet submitted for review. It covers one harness and one environment, the temporal result rests on a pre-written permission rule and a synthetic session ID, and we found no released code. We read the full text.

**Why it matters:** An approval that records a command string says nothing about what hooks or scripts that command triggers. If you rely on allow rules in a repository with commit hooks, review what those hooks do. Whether these rates hold outside the author's setup is untested.
