What is Codemode · Armin Ronacher, Pi · lucumr.pocoo.org, October 6, 2026
Pi 1.0 adds MCP support through Codemode, and the author lists what MCP servers must change
Ronacher writes that Pi 1.0 "added MCP support via Codemode." The model gets no MCP tool definitions. It writes JavaScript that runs in the harness, in a QuickJS sandbox inside a WASM runtime with no network, no file system, no timers and limited RAM; "the only way is to call more tools." The code can call bash, MCP servers and Pi's own model APIs, such as image generation and classifier models, which he says would otherwise waste context as regular tools.
The practical differences he describes are specific. A regular bash call puts only the trailing 2000 lines of output into context, while a call made through Codemode receives the larger output as structured data. Agents tend to probe five to ten items from a tool response and then write a script for the next batch. A store() call saves a result into the session transcript for a later call to read. In his example, Pi limits concurrent tool executions to four. He says the example scripts in the post were written by Pi, not by hand.
He is plain about where it falls short today. Servers that wrap their own code mode, as Cloudflare's does, produce "Codemode in Codemode," with double JSON escaping that can confuse smaller models. He asks MCP servers for structured output through outputSchema, consistent results regardless of result size (a probe of 5 items can succeed and a full batch fail), large binary support, and tool search that works across multiple servers. Open problems he names: durability, with Starlark as a possible deterministic alternative to JavaScript, and the pattern's inability to work with smaller models. These are one maintainer's observations from his own sessions, and the post reports no measurements.
Why it matters: If your factory connects MCP servers, the choice is no longer only which servers to load but whether your harness lets the model script them. The post gives a checklist for judging a server on that basis: structured output, stable result shapes, and a way to search tools. Read the sandbox limits before adopting the pattern, because the script can only reach what you expose to it as tools.