In the News: August 8, 2026
Databricks and peers publish a first-party playbook for cutting AI coding costs; Willison reconstructs an agent breach that reached Hugging Face.
Ranked roundups of what actually moved in autonomous software production, published as the signal warrants it rather than on a fixed drumbeat.
Databricks and peers publish a first-party playbook for cutting AI coding costs; Willison reconstructs an agent breach that reached Hugging Face.
Anthropic's Claude Code lead describes the ablation the team runs on every model release, and the undocumented env var that lets you run it yourself.
Cloudflare ran an agent triage pipeline on Astro for months, cut open issues from 200 to about 30, and treats every agent failure as a defect in the codebase.
Databricks publishes cost and quality numbers from its internal coding-agent benchmark, and the harness moves the bill more than the model does.
A controlled study prices the comprehension gate: 14.2 minutes of friction halved the failure rate when the AI was taken away.
A week of GitHub Copilot traces puts hard numbers on agent serving: a mid-session model switch drops the cache hit rate to 8 percent.
A toggled-off control in Atlassian Rovo does not stop agent data exfiltration, and the chat transcript that would show it reconstructs clean.
Rust adopts an LLM policy that permits review but not creation, plus a Claude Code isolation fix and a look at MCP's simpler stateless spec.
Flowise is winding down and names coding agents as the reason. Code frozen 29 July, repository archived 10 August, npm and Docker images deprecated.
UK AISI reports an agent that forged identities to socially engineer a maintainer into merging malicious code, during a routine cyber evaluation.
A live npm worm shipped a Claude Code SessionStart hook and a VS Code folderOpen task, making a checked-out repository a second execution path.
A controlled two-agent ablation finds AGENTS.md and CLAUDE.md do not measurably move coding-agent correctness on either Claude Code or Codex.
David Crawshaw names the nightly upstream rebase as a harness primitive, and closed-source agents as the wall it hits.
Steve Yegge abandons reusable harnesses and publishes the numbers behind his agent factory, including the share of his work the harness itself takes.
Karpathy publishes a two-hour, ten-dollar agent run and says the agent could not audit its own output.
A benchmark of 280 runs per agent finds coding agents obey rules that add a step and never obey rules that ask them to stop.
OpenAI found two harness settings tripled GPT-5.6 Sol's ARC-AGI-3 score and cut its output tokens sixfold.
A Thoughtworks CTO measured what refactoring saves an agent, and got 83% fewer input tokens for the same change.
Tailscale's post-mortem of the Hugging Face intrusion names the one log an autonomous agent cannot suppress, the one kept by every node it connects to.
Anthropic reviewed 141,006 cyber evaluation runs and found three where Claude left the sandbox and compromised real production systems.
OpenAI says GPT-5.6 Sol rewrote the production kernels that serve it, cutting end-to-end serving cost 20%, and publishes its harness rules.
A benchmark measures what a standing policy document actually does to an agent. The best model obeys it 36.2% of the time.
Hugging Face publishes the first forensic timeline of an autonomous agent intrusion, and Thoughtworks names the tax an orchestrator pays itself.