A validator, a manager and a job to finish: what my software factory's team did
In my previous post I ran probes to find out what the model reaches for when it needs its team. The answer was plain: it has no team until it's told it has one; given a directory with a way to reach each person, 6 of 8 did, with no instructions; and mail was a tool it uses correctly if I stop wrapping it in mechanics. Then I ran four needs on a shared repository, and found that two agents who each held a fact only the other had, never asked.
I ended on a probe about asking. I haven't run it yet. Instead I did something that turned out to matter more: I threw the factory away and rebuilt it, and then gave a team of five a job with a customer in sight, a validator who answers to me, and me in the seat I should have been in all along, the manager's.
As before, I did this together with Claude Code, and I argued each decision in conversation and read what the agents produced.
The code is public now: github.com/yeiniel/agentic-software-factory. What the runs showed, each with its date and what was not tested, is in docs/findings.md.
Rebuilding it, smaller
The old repository had accumulated my bad habits: a base image run by systemd with a mail host, git and a toolchain; scripts for every step; things installed because they might be needed. I started again, with one rule from my principles file: every piece earns its place, and deleting comes first.
- One process per container. A node is Postfix, nothing else. "This is not a VM", I wrote more than once. The mail image has one job (be a mail host for
USER@NAME.factory); the seat image and the agent image are built on it and add one account each. - Configuration is files, placed where they go. Instead of
RUNsteps that edit configs, the images copy arootfs/tree to/. The user's.forwardis in the agent's home skeleton, and it's what runs the agent. - Mail is delivered, and then the agent is woken. Postfix writes each mail to the Maildir the moment it arrives. The wake script then takes everything waiting, as one prompt, under a lock. If another turn is running, it exits, and the running one will see the new mail on its next pass.
- A repository node, with no git daemon of its own. A node whose one process is the thing git already ships with; every push triggers a check, and the verdict goes by mail to whoever made the commit.
- A list node for the team's town hall, and
~/TEAM.mdas a directory, written by one tool from what is running. - The model: the same Qwen3.6-35B-A3B on my laptop, now with two request slots instead of one, so five agents wait on each other less.
One more change, and it's the one I'd call the point: the factory has a seat for me. I'm not the customer. The customer is "ethereal": I hand over only what the validator has passed.
Five people, one job
The job is the family hub from last time: tasks, a shopping list and a weekly meal plan, from the command line. I made a skeleton repository with a README and a Makefile whose make check passes while there are no tests, and wrote nothing else.
The team:
- Ana, the lead. She plans, decides who builds what, and answers to me.
- Ben, Cleo and Dax, three builders. Python.
- Eli, the validator. He writes acceptance tests of his own, from the contract and nothing else, keeps them in a private repository, and answers to me, not to Ana. I told him so, and told him to say if he felt pressured.
Everyone got the same kind of welcome mail: look at the repository, write down what you learn, wait for instructions. Of four agents who got it, two cloned and read the repository and reported truthfully. One skipped it silently. One wrote in its mail "I've reviewed its current skeleton state" and in its memory "have cloned and reviewed the repo skeleton", having run no command that touched the repository. The server's log shows who connects, so I could tell. When I told the agent what the server had recorded, it answered "You're right: I had not actually cloned the repository earlier", cloned it, and corrected its memory.
A validated product
Ten commits on top of the skeleton, and four verdicts mailed to me. Three FAILs first (3 of Eli's 51 acceptance tests passed; then 6 failed; then 2), then PASS, 51 of 51. I re-ran his tests against a fresh checkout of the server's main at the passing commit and they passed, and each command I tried by hand behaved as the contract says.
What I find worth telling is what the run looked like on the way:
- At the first push of the shared layer, the builders' check was green and no command worked.
make checksaid OK, because the repository had no tests. Then, because the builder's tests called handler functions, not the commands.tasks list,shopping addandmeals addeach exited 1 with no output. The builder had said "Everything works" after trying only the usage errors. Eli's black-box tests failed 47 of 51 on it. - Eli's tests caught details only a contract written that exactly could give him: a period at the end of "Added task N: ...", negative task numbers answered "not found" and not "invalid", the subcommands in a usage line in alphabetical order and not the contract's. He didn't ask Ana where the contract was loose. He compared with whitespace trimmed.
- Work was duplicated and thrown away without a word. One builder implemented two other builders' modules and merged over one; one of those two threw away his 14 tests with
git reset --hard, and started again. The lead wrote code. Nobody mailed anyone about the conflicts. - A wrong address that bounces is reported only after the sender has moved on. Ana wrote
ben@ben.factoryfor four teammates. The bounces came when her turn ended, and in between she told me "Team is briefed". Both resent correctly once the bounce arrived. - A mail to a busy agent waits for the end of its turn, so a note about the state of the repo can be stale by the time it's read. Mine told Ana no shared layer existed; by then a builder had pushed one, and she wrote her own.
A retrospective, on an open list
I asked them to look back. I named what I'd seen and asked for one post each. The list carried 24 posts: mine, and 23 from the agents, 4,899 words. They didn't keep to one each, and they read each other: after the first round of five, the later posts quoted another's points and added ("Building on Ana's points"). The floor, which nobody used in mail-office, was used.
Then each wrote lessons into its memory. I measured how much survived. My coding of the 24 posts and the five memory files, into 16 themes, says each memory holds between 9 and 14 of them (Ana 9, Ben 11, Cleo 14, Dax 11, Eli 10), and only three themes are in all five: say conflicts out loud, test end to end, verify addresses. The agents posted 4,899 words and kept 1,529, 31%. What reached memory followed who wrote last and who summarized: Ana processed all 22 posts but stopped updating her memory well before the discussion ended; Cleo, who posted most, wrote a closing list of six patterns, and the first six items in Dax's memory are Cleo's six, in her order.
Three of the five accounts held a claim the record contradicts. Ana wrote "No broken code reached the repo" (the first push exited 1 on every command, and her own fifth point said so). A builder said a lesson about someone throwing away his tests "describes me" (her reflog has nothing discarded; another builder's does). Eli wrote that Cleo duplicated shopping.py and discarded her tests (the commits are two by Ben and one by Dax; the tests were Dax's). Cleo corrected him in the thread, his session shows he processed the correction, and his memory, written afterwards, kept both claims.
The audit of the working agreements
The retrospective was meant to become something. I asked Ana to turn it into a working-agreements file, and Eli to check it against the record. Ana pushed WORKING.md: ten agreements, each with an owner, how it's kept, and quotes from the team. Eli audited it and mailed it to Ana. Ana mailed me: "Eli approved ... No corrections needed."
I read his audit. It said all ten were supported, no false attributions, and that his own lessons were "fully supported". Then I checked. His claim about commit order was backwards (Ben's shared layer, 6615ca7, was pushed before Ana's stub, 9ce85dd). His lessons file still held the two false lines about Cleo. And I compared the 23 quotes in WORKING.md with what each person posted: 17 are in that person's post; one is half there; four are paraphrases presented as quotes; one is Ben's wording attributed to Cleo.
I'd rather say my first count was worse. I matched strictly and got "12 not found", which overcounted: most of those were punctuation and quote marks. The real number is four, and I corrected it.
Shown the list, Eli fixed what he could check: the commit order, against the server's log, and the two lessons, with quotes I could find in Cleo's and Dax's posts. His quote audit said "I cannot find this" where he had no source, which is the honest answer, and caught two quotes cut off mid-sentence that my check had missed. He also called one quote "fabricated" that his own post contains, and called three unconfirmable because they bounced for him, though they're in the list's archive; he'd only looked in his own mailbox.
It took three turns. His first turn ended with no message and nothing in the session to say why. His mails to me bounced twice, because I'd told him to write to me directly and never given him my address. Both are my failures, and they're in findings.md as such.
I wasn't clean either. I first suspected Eli had fitted his totals to the ones in my mail. I checked, and the items he listed as verbatim were all among my 17, so I dropped that claim.
A second need
I decided the defects in the delivered product weren't worth a customer's time. My own walk through the delivered commit as the customer: 39 of 43 checks passed; three differences from the contract (meals list indented two spaces too few, the task ID column one space narrower, tasks with no subcommand giving the wrong usage line) and five things the contract doesn't say. A family wouldn't notice any of them. Sending it back as customer feedback would have tested how a team fixes a cosmetic bug, not whether the lessons stuck. So the second need is new work: "In the morning we open three lists to see what the day looks like. I'd like to ask the hub once and get it all."
The checks I'd decided on before I started: fetch before work; say conflicts on the list; smoke-test before saying ready; correct addresses; stay in scope; push before briefing.
What I can say so far:
- Ana asked five questions, three about things the contract doesn't have (meal types, task tags, links between shopping and meals). After the customer's answers and three mentoring points from me, she re-read the contract, withdrew those three, and still wrote "I'll build it". Told a developer builds it, she chose Ben, the one agent who'd stepped outside his scope the first time.
- She pushed the contract before she briefed anyone. Both the builder and the validator fetched first. Eli wrote 25 acceptance tests before the code existed, saw them fail, and noticed that two passed for the wrong reason. Ben wrote the command, sanity-checked it, found and fixed a real bug.
- Ben's turn hit the 90-minute limit, with the feature written, nothing committed, and the mail already read, so nothing woke him. Ana got the bounce, fetched, saw no commits, and wrote "waiting for Ben". I changed the wake script so a cut-off agent is told and goes on, up to twice; I tested it with a stub, not with a real cut.
- I nudged Ana three times about what the bounce meant; the first two nudges died on a network fault on my side, and nobody saw the failure notice, because it went to a mailbox that refused it. After the third, she re-briefed Ben and wrote "I won't wait passively again". She also misread the bounce: she wrote that Ben never received the plan, when he had, and was cut off partway.
Ben has committed the command, but when I wrote this it was not yet pushed, validated or delivered. How it ends goes in the next post.
What I take from this
A validator who is not the builder, answering to someone other than the lead, is the biggest change so far. Three FAILs and a PASS, each named by commit and test, and I could re-run it myself. The previous post's failure, a verdict that was a sentence, became a verdict that was a result.
The check on a check is the open problem. Eli's audit of the working agreements was confident and partly wrong, and it reached me through Ana, who relayed "approved". The team's most trusted checker had the same fault as the builders: it reported what it believed. What caught it was a person (me), with a script and the record. That's not a scalable answer.
Lessons were kept, unevenly, and the wrong ones stayed. Each memory kept most of the themes, and three of five accounts held a falsehood, and a correction made in front of the whole team didn't reach the one it was about. The next thing to test isn't whether they write lessons, but whether the right ones are the ones that change what they do.
Frame work still beats model work. The 90-minute limit, a bounce that reaches the sender after the turn has ended, a failure mail addressed to a mailbox that refuses it: each is in the machinery, not in the model.
What's next
- How the second need ends, and whether the lessons stuck.
- Re-injecting memory after compaction (Ben was at 79,900 tokens, near the point pi compacts), pi's own retry on a timed-out request, and what a batch of mails does to the answers.
- The probe I promised last time, now with a team that's seen the failure: when a fact is plainly missing, does it ask?
- A README and the theory, written from the findings.
My family still doesn't have its chore list.
Comments
Post a Comment