Building a software factory out of Unix mail

"Agentic software factory" became the phrase of 2026. StrongDM decided no human would write or review code. OpenAI shipped about a million lines in five months with a handful of engineers. Uber says about 70% of its pull requests now come from agents. The picture is always a production line: planner agents, coder agents, tester agents, reviewer agents, and engineers who design the line instead of writing code.

I wanted to try one. Before building anything, though, I asked a more basic question: why would I want a factory at all, rather than a single actor that produces software? This post covers where that question took me, from first principles to two small proofs of concept built with the most old-fashioned tool I could find: Unix mail.

I did this exploration together with an AI coding agent (Claude Code), in a loop of research, argument and building. Some of the best turns came from pushing back on its answers, and I'll point those out as they come.

Does software production scale with more actors?

Before copying a factory, I wanted to know whether adding actors to software production helps at all. It turns out to be well studied, and the answer is only partly.

Neil Gunther's Universal Scalability Law models it with two costs:

  • Contention: work that must happen in sequence, such as one shared spec or one reviewer. This alone gives diminishing returns (Amdahl's law).
  • Coherency: the effort of keeping everyone's understanding in sync. The number of pairs grows as N², so past some point adding actors makes the team slower. Brooks's law ("adding people to a late project makes it later") is this term in practice.

The data on human teams agrees:

  • QSM: large teams bought only about 24–30% schedule compression, at 3–4× the cost and 2–3× the defects.
  • Apache httpd: about 15 core developers wrote about 80% of the code.
  • DORA: teams grow without slowing down only when the architecture is loosely coupled. Scale is bought with modularity, not headcount.

The deepest explanation I found is Peter Naur's 1985 essay Programming as Theory Building. A program isn't really its code. It's a theory in its builders' heads about why the code is the way it is, and code and documents record that theory only in part. Every new actor has to rebuild the theory, or the program degrades.

Early research on agents points the same way. A Google/MIT study found multi-agent setups lowered performance by 39–70% on sequential tasks, and raised it only on work that splits into independent pieces. Another study traced about 42% of multi-agent failures to bad specifications and 37% to agents misaligning with each other. These are the same failures human organizations have.

So more actors help with one thing, the throughput of independent work, and usually hurt speed, quality, leanness and the ability to change course.

Specialization is an inherited paradigm

This is where I pushed back. Much of what we "know" about team size and organization comes from a world where the actors were humans, and humans specialize. Agents don't need to. Shouldn't that change the picture?

Partly. Human specialization answered two limits:

  1. Skill acquisition. Learning a skill takes a human years, so we split work by skill, and got handoffs, layered teams, queues between specialists and a lot of the N² coordination. Agents remove this limit. One agent can write the migration, the API, the UI, the tests and the deploy config in one line of reasoning, with no handoff.
  2. Cognitive capacity. One mind can hold only so much. Agents keep this limit, in a new form: the context window, loss of quality in long contexts, and memory that resets every session.

Two consequences follow. First, the natural unit of production is one agent owning a vertical slice end to end, sized to what its context can hold and reload. The classic planner → coder → tester → reviewer pipeline of agents rebuilds exactly the handoffs agents make unnecessary. It's a human pattern copied onto agents.

Second, some separation is still worth having, but for independence, not skill. A verifier with a fresh context doesn't know more than the builder. It just doesn't share the builder's assumptions. The one specialization left belongs to humans: owning intent, knowing what the product is for and what "good" means.

What a factory actually is

My first attempt at a definition was "a collection of agents collaborating on a piece of software, plus the tooling for their communication". Whether they own verticals or slice by layers is then just configuration for productivity.

The agent's first answer was that collaboration is optional, because a factory with one actor is still a factory. I disagreed: even with one actor, you still have to communicate intent to it. That turned out to be the key point. Communication isn't a feature of multi-agent systems. It's the substance of the factory, even with a single agent. The definition we settled on:

An agentic software factory is a set of actors (agents and humans) connected by communication channels, arranged in a repeatable loop that turns intent into verified software.

A one-actor factory already has five channels:

ChannelDirectionCarries
Intenthuman → actorthe goal and what "done" means
Report / escalationactor → humanresults, evidence, questions, "I'm stuck"
Realityactor ↔ workspaceactions and what actually happened
Verdictverifier → actorpass/fail and why
Memoryactor now → actor laterthe product's theory, decisions, state

A sixth channel, peer, appears only with more than one actor. The memory channel is the most important one. Sessions reset, so one agent is really a series of actors across time: every new session is a newcomer rebuilding Naur's theory from whatever the last one wrote down. The coordination cost of teams shows up even at N = 1, sequentially instead of in parallel.

That gives the minimal factory: one builder, an independent verifier and a human, connected by those channels, running in a loop. It has three tests: a new intent goes in and verified software comes out without touching the process; the next job goes better because of what the last one wrote to memory; and the builder cannot declare its own work done.

Did OpenClaw reinvent the wheel?

The obvious reference for "one agent I can talk to in real time" is OpenClaw: a gateway to chat apps, an agent loop, a workspace of Markdown files, memory, a heartbeat. Mapped onto the channels, it covers intent, report, reality and memory. It has no verifier and no acceptance gate. It's a minimal actor with a communication shell, not a minimal factory. It's also about 430k lines, which conflicts with two principles I wrote down for this project: own the thing (small, self-hosted, auditable parts) and no vendor lock-in (model, transport and storage behind thin interfaces I control).

Then I remembered logging into a Linux box years ago and seeing:

You have new mail in /var/mail/yeiniel

Unix has had most of this plumbing for decades:

  • Mail for durable, asynchronous messages between users.
  • Maildir (tmp/ → new/ → cur/) as a crash-safe queue where claiming a message is an atomic rename.
  • biff and comsat to announce new mail on your terminal.
  • systemd .path units to start something the moment a directory stops being empty.
  • Unix users as trust boundaries between actors.

A lot of what the claw projects build (queues, dispatch, schedulers, logs) is a reimplementation of this. What Unix doesn't give you is a way to reach a human who isn't logged in. Email, though, is an open, federated, self-hostable transport, which fits better than any chat API.

POC 1: an agent that is just a Unix user who gets mail

The first experiment was deliberately naive. One container runs a mail server (OpenSMTPD) with two users, human and agent. When mail lands in the agent's Maildir, a small biff script starts pi, a minimal open-source coding agent, on a model running locally in LM Studio. The whole prompt is the line Unix used to print:

You have new mail in /var/mail/agent

It worked. With no framework and no tool definitions for mail, the model listed the Maildir, read my message and replied. My setup was broken at first (Debian's package doesn't install a mail command), and the agent searched for another client, found s-nail, sent the reply, and still followed its instructions to mark the message read.

It was also slow and expensive. A "nice to meet you" mail took 14 model calls and about 36,000 input tokens. The main lesson from POC 1 was about cost: every call resends the whole context, so cost is roughly calls × context size. Only three of the 14 calls did real work. The rest went to finding the mail, doubting the mail was complete, creating a folder and checking its own steps.

The fix was the Unix one: put work that is always the same into a tool, not the model.

  • biff started announcing each message the way comsat did, with the file, sender, subject and first lines, ending with [end of message] so the model stopped wondering whether it was truncated.
  • A one-line reply <file> command took care of the subject, recipient, sending and marking the message read.

The same mail went from 14 calls and 36k tokens to 3 calls and about 6k tokens.

POC 2: a builder, a verifier and a post office

Several things in POC 1 bothered me. The mail server lived inside the agent's container. biff was a shell loop that woke up every minute to check, which is wasteful. The agent was hard-coded. The second POC fixes those and takes a real step toward the minimal factory:

  • The post office is its own container. OpenSMTPD handles authenticated submission and Dovecot serves IMAP. Every actor logs in with its own credentials and can send only as its own address.
  • Actors are provisioned, not hard-coded. ./provision builder creates a folder with the actor's identity, mail credentials, server and model settings, plus its instructions. That folder is mounted read-only into the actor's container. The same image runs twice, once as builder and once as verifier.
  • No loop of my own. systemd runs inside each actor. fetchmail waits on IMAP IDLE, so the server pushes new mail, and files it into ~/Mail/new. A .path unit with DirectoryNotEmpty= starts the agent. Nothing polls.
  • A mail thread is an agent session. The session id is derived from the thread's first Message-ID. When the verifier's verdict arrives in a job's thread, the builder continues the same session, and remembers the job, instead of starting from nothing. This is a cheap, standard answer to part of the memory channel: References: headers have threaded conversations since the 1980s.
  • Filesystem conventions as a skill. Instead of a hard-coded ~/work, agents get a small skill describing the XDG user directories. xdg-user-dirs now includes ~/Projects, so that's where projects go.

The builder's instructions encode the most important rule of the minimal factory: you cannot declare your own work done. It does the job, mails the verifier the job as I asked it together with the evidence, and replies to me only after a PASS. The verifier checks work it didn't do, answers PASS or FAIL in the thread, and escalates to me if the job itself looks wrong.

All the channels map onto mail: intent is my mail to the builder, verdict is the verifier's reply, report/escalation is a reply to me, and memory (short-term, for now) is the session attached to a thread. The human side is a few lines of curl over SMTP and IMAP, or any mail client.

What happened when I ran it

The plumbing worked on the first real run. I mailed the builder a Node.js fizzbuzz. It read the filesystem skill, created ~/Projects/fizzbuzz, wrote and ran the code, and mailed the verifier. It did not reply to me, just as its role says. The mail landed in the verifier's box through IMAP IDLE and systemd started the verifier in the same thread.

Three failures turned out to be the most useful results.

The first: my first test job asked for fizzbuzz in Python, and the actor image is built on Node and has no Python. The builder spent twenty calls looking for an interpreter that wasn't there. The environment is part of the intent: an agent can't satisfy a job its workspace can't run.

The second: the verifier went looking for the builder's files on its own machine, didn't find them, and started to conclude that the evidence was false. In its own reasoning it even considered that it might be on a different machine, and then dropped the idea. This is the missing reality channel showing up: the two actors share mail but not a workspace, so the verifier can only check what the builder chooses to put in a message. Verification is only as independent as the builder's quote of the job. That's the gap the next POC has to close.

The third: after 26 minutes the verifier did reply, and the threading worked. The verdict arrived with the right References: header, and the builder picked up its original session and remembered the job. But the verdict was empty. The verifier had called reply without passing a body, the agent's shell gives commands an empty stdin, and my tool sent a blank mail without complaint. The builder then spent more than twenty calls and about 150,000 input tokens taking that mail apart with cat -A, xxd, od and strings, looking for a PASS or FAIL that wasn't there.

The fixes were small. The verifier's instructions now say plainly that the builder's files are not on its machine. send-mail now refuses an empty body and tells the agent how to pass one, so the mistake costs one call instead of a hundred thousand tokens in someone else's session. With both in place, the rerun is next.

What I take from this so far

  • A factory is its channels. Once I stopped thinking about roles and started thinking about intent, report, reality, verdict and memory, every design question became concrete: what does this channel carry, who can write to it, and how long does it last?
  • Old primitives go a long way. Unix users, mail, Maildir, threads and systemd units cover identity, queues, activation and conversation history, in standard, auditable pieces I can run myself. The agent itself is a small program that reacts to "You have new mail".
  • Token cost is a design problem. Every deterministic step the model does costs a call times the whole context. Pushing those steps into small tools cut cost by about 6×, and made the agent more reliable.
  • Tools must fail loudly. An agent believes what its tools tell it. A tool that quietly succeeds at the wrong thing moves the error into another actor's context, where it is far more expensive to find.
  • Verification needs shared reality. Separating the builder and the verifier is easy. Giving the verifier something real to check, with criteria the builder can't change, is the hard part. That matches what the industry reports: generation is cheap, verification is the bottleneck.

What's next

The next POC adds a shared workspace, so the verifier checks the real artifact instead of a description of it. I also want the human's acceptance criteria to reach the verifier directly, not only through the builder. After that come a lasting theory/memory that survives beyond a single thread, and an acceptance gate that decides what joins the product. Then I can check the three questions: does a new intent flow through without touching the process, does the next job go better because of the last one, and is the builder really unable to declare itself done?

I'll write about it when I get there.

Comments

Popular posts from this blog

Concurrent Alarms Processing

Creating a package for Debian, with its headaches

AI al inicio de 2026