---
title: 'Portfolio, demand & product discovery'
description: 'Deciding what is worth building before the factory optimizes how to build it.'
canonical_url: 'https://darkfactory.dev/factory/portfolio-demand-discovery'
markdown_url: 'https://darkfactory.dev/factory/portfolio-demand-discovery.md'
collection: factory
date_published: '2026-07-16T00:00:00-04:00'
date_modified: '2026-08-09T00:00:00-04:00'
---

# Portfolio, demand & product discovery


**Confidence: low.** *Evidence: theory and adjacent case studies.* *Last substantive change: 2026-08.*

Before a factory optimizes how it builds, it has to decide what is worth building at all. Demand covers where work comes from, which ideas get funded, and whether a metric moved because customers received lasting value.

## The conclusion

**Product judgment sits upstream of almost everything the published evidence describes, and it remains human.** Nearly every documented factory starts after someone has already chosen the work and prepared an issue queue. As the cost of implementation falls, the quality of demand becomes the binding constraint: a factory that builds the wrong thing faster is a faster way to be wrong.

## How the thinking got here

Early accounts began with a chosen task or a groomed backlog and asked only how fast agents could clear it. Cloudflare and Astro's [open-source maintenance factory](https://blog.cloudflare.com/astro-issue-triage/) extends automation upstream into issue reproduction, diagnosis, and preview releases, but its demand still arrives as human-filed issues; it does not discover which product outcomes deserve investment. As throughput rose, attention moved upstream. If output is abundant, the scarce input is a good decision about what output is worth producing, and non-software autonomous-business experiments have repeatedly shown agents optimizing a proxy metric straight into Goodhart failures.

## Credible alternatives, and when each is right

| Approach | Right when |
|---|---|
| Human-owned portfolio | high-stakes bets; strong domain judgment available |
| AI-assisted discovery | large opportunity space to survey; humans still fund |
| Agent-proposed experiments with human funding gates | cheap, reversible probes with clear kill criteria |
| Autonomous optimization against an explicit metric | the metric is trustworthy and hard to game |

## Where it fails and what we still don't know

The failure mode is Goodhart: an agent moves a chosen number without moving real value. The open questions are large. How does a factory discover unmet needs, kill weak ideas early, incorporate user research, reconcile conflicting stakeholders, and separate metric movement from durable customer value? The evidence here is the thinnest in the whole map.

## What would change our mind

A documented software factory that discovers and validates demand autonomously, with durable customer outcomes rather than metric movement, would move this from human-owned to genuinely shared. None exists yet.

## Evidence and further reading

- [Bengt Hires a Human](https://andonlabs.com/blog/bengt-hires-a-human)
- [Vending-Bench](https://andonlabs.com/blog/openai-gpt-5-5-vending-bench)
- [Autoresearch as a Production Loop](https://shopify.engineering/autoresearch)
- [How we built a software factory to drive Astro's GitHub issue count to zero](https://blog.cloudflare.com/astro-issue-triage/)
