← In the News

Stripe moved factory orchestration out of the LLM and into code

Stripe's Payment Method Factory: Orchestrating agents for repeated, custom integrations · David Dunne, Xenofon Vourliotis and Sai Samant, Stripe · stripe.dev, September 30, 2026

Machine-readable Download Markdown

Stripe says the factory has created three new payment method integrations and migrated ten existing ones, cutting integration timelines "from as long as six months to two-to-six weeks." These are Stripe's own figures. In one bake-off, two engineers did the same integration task: one used vanilla Claude Code and needed roughly 20 days, the other used Claude Code with a reusable prompt set and needed 4. That is one task and two engineers.

The design has three parts. An observer agent reads each implementation run's transcript and reports where the agent got lost or had to re-learn something, and an engineer uses that report to refine the prompt. After each integration, an agent reads the learning logs, decision logs, transcripts and review comments and proposes prompt updates, some of which apply with one click through Stripe's Minions coding agents. And the orchestration layer changed: the first orchestrator agent "worked, but it was expensive, slow, and sometimes unreliable," and sometimes tried to do a step itself instead of delegating it, which used up its context. Stripe's conclusion: "predictable coordination is better implemented in code."

The scaffolding step shows the effect on a single task. Stripe says it was a change of more than 5,000 lines across over 100 files, driven by 15 commands that must run in a set order, and a prompt explaining each piece cut it from up to a week to a day. The post also says the factory snapshots a provider's sandbox responses into immutable contracts instead of hand-written mocks, and that Stripe's tax and verification teams built their own factories after seeing this one.

Why it matters: If you run agents over a repeated task, the post is an argument for putting dispatch and sequencing in deterministic code and keeping the model for the steps themselves. It also describes a maintenance loop worth copying: every run leaves logs, and an agent turns them into prompt edits that a person approves. The numbers come from Stripe and the bake-off is n=1, so treat them as direction, not a benchmark.