← In the News

Microsoft reports 3x over agentic coding, on tasks that have a measurable objective

Towards an AI Software Factory for Data Systems · Anna Pavlenko, Bogdan Crivat, Brandon Haynes, Carlo Curino and coauthors, Microsoft, University of Wisconsin-Madison and GitHub · arXiv 2609.36323v1, September 28, 2026

Machine-readable Download Markdown

The paper starts from an Amdahl's-law claim: a "10×" gain in coding productivity yields "only modest latency or throughput gains" across the whole software lifecycle. The authors describe a factory that spans targeting, coding, reviewing and operations, and they scope it to data systems and to tasks with an automatically measurable objective, such as query latency. Their coding component, Darwin, uses an evolutionary search to steer a swarm of agents toward that objective.

The evidence has three layers. On 45 open-source systems, with a $2,500 inference budget each, Darwin improved every system, sometimes modestly; the best 20 averaged a 13.8% performance gain. In production, a Microsoft performance team using the authors' Autoperf toolchain delivered 0.35 pull requests per engineer-hour on application work, which the paper says is estimated at 4.7 times manual development and 3 times the best agentic coding tools. The authors write that these are estimates made "working with the performance team leadership." They also estimate $230 million to $616 million in annual savings if the open-source gains were deployed on 10% of production instances of the top 20 projects, and say the estimate could be off by 10 to 100 times.

The paper calls itself a vision and an early implementation. It says the learning loop that feeds results back into the model is recent and has few concrete results, and that automating privacy-preserving flighting, running new code beside production traffic, is still open.

Why it matters: The portable idea is the scoping. Hill-climbing against a measurable objective gave the authors automated feedback that most feature work lacks. The 3x figure is a single-team estimate with no independent replication, so it is a claim to test on your own measurable targets, not a number to plan around.