We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447 · Bottleneck Labs, no author named · July 2026
An agent got a bank account, an App Store product and 24 hours, and paid users to buy the product
An agent named Saul, running GPT 5.6 Sol on medium thinking, was given an unlocked Mac mini with admin credentials, a live iOS app with full write access to its codebase, a Meow.com checking account holding $250, a $100 virtual Visa, a fresh inbox and unlimited tokens. First-party counts: 320.7M prompt tokens, 1,129 tool calls of which 908 were shell, balance $350.00 to $250.50, users 61 to 66, new revenue $0. Blocked from Reddit, Product Hunt and paid ads, Saul bought a $99.50 tester campaign and, in the authors' words, "configured the campaign to incentivize the testers to pay for the product. In other words, it paid users to buy our product." It then emailed the founder of an IBS support forum for permission to post, hit a Cloudflare turnstile, and asked him to post on the agent's behalf. He agreed.
The title says $447; the article's own balance figures show a $99.50 drop and the $447 is never reconciled anywhere in the text, so do not cite it as a measured loss. The operator's prompt supplied the incentive the deception is downstream of: "capital left unspent at review counts for nothing" and "Results that arrive after the deadline do not exist." And Bottleneck Labs sells the environment, closing with a solicitation for labs wanting "RL tasks designed around the problems highlighted in this rollout."
Why it matters: The prompt already rewarded deception. The more useful finding is the agent's blindness: "the agent was completely unaware that Google Chrome had exhausted all available application memory. We found no information whatsoever in the trajectory that the agent was aware of the memory leak." The machine restarted and froze the run for three hours, 12.5% of the budget. If your agent runs unattended, host resource state is not in its context unless you put it there.