---
title: 'In the News: September 26, 2026 (Evening)'
description: 'A small open-source project pairs a coding agent with an independent supervisor process, and says openly that its own judgment is not yet proven.'
canonical_url: 'https://darkfactory.dev/news/2026-09-26-evening'
markdown_url: 'https://darkfactory.dev/news/2026-09-26-evening.md'
collection: news
date_published: '2026-09-26T19:20:00-04:00'
date_modified: '2026-09-26T19:20:00-04:00'
---

# In the News: September 26, 2026 (Evening)


Tonight's edition is a single item: a small, open-source project that pairs a coding agent with an independent process built to judge its progress, and that says plainly it does not yet know if that judgment is any good.

## 1. An open-source "foreman" watches the coding agent instead of writing the code

**[Foreman](https://github.com/thruwire/foreman)** · thruwire, GitHub · source repository, published September 2026

Foreman is a Python runtime that pairs a Codex coding-agent worker with a second, independent process built on TypeSafe AI's Jev model. While the worker reasons, edits, and tests, Foreman watches the same repository state and answers nine yes-or-no probability questions about it, including whether the worker is stuck, whether requirements are satisfied, and whether the job is ready to finish. A deterministic Python policy, not the model itself, turns those answers into one of seven actions: continue, start a worker, start a verifier, stop, retry, finish, or escalate to a person. The project ships a full offline test suite and a credential-free demo, but its own documentation says "Jev assessment accuracy is unproven for this use case," and the author calls the whole thing an experiment, not a finished alternative to existing harnesses.

**Why it matters:** An independent process that judges a coding agent's progress from outside its own reasoning loop has mostly been an idea discussed in the abstract. Foreman is a small, working instance of it that anyone can install and read end to end, and the limitation it admits up front, that the supervisor's own judgment is not yet calibrated, is the question worth tracking as more of these appear.
