← In the News

An open-source "foreman" watches the coding agent instead of writing the code

Foreman · thruwire, GitHub · source repository, published September 2026

Machine-readable Download Markdown

Foreman is a Python runtime that pairs a Codex coding-agent worker with a second, independent process built on TypeSafe AI's Jev model. While the worker reasons, edits, and tests, Foreman watches the same repository state and answers nine yes-or-no probability questions about it, including whether the worker is stuck, whether requirements are satisfied, and whether the job is ready to finish. A deterministic Python policy, not the model itself, turns those answers into one of seven actions: continue, start a worker, start a verifier, stop, retry, finish, or escalate to a person. The project ships a full offline test suite and a credential-free demo, but its own documentation says "Jev assessment accuracy is unproven for this use case," and the author calls the whole thing an experiment, not a finished alternative to existing harnesses.

Why it matters: An independent process that judges a coding agent's progress from outside its own reasoning loop has mostly been an idea discussed in the abstract. Foreman is a small, working instance of it that anyone can install and read end to end, and the limitation it admits up front, that the supervisor's own judgment is not yet calibrated, is the question worth tracking as more of these appear.