Back in the Saddle
I ran boards for human teams for years. Now I run one again, and every worker is an AI agent. How I built Software Factory, what went wrong first, and what changed in how I deliver.
◆ marks the two gates I own: approving a design, and promoting Final Review to Done.
Two agent platforms, one tired manager
A few weeks ago I was running Codex and Claude sessions most of the day, several of each at once. Every one of them was a capable worker. Together they were a mess, with the same problems every manager knows, only faster.
Context switching
blockerEvery session had its own tab, its own history and its own half-finished idea of the plan.
No shared memory
blockerA decision made with Claude in the morning meant nothing to Codex in the afternoon. I was the only link between them.
No paper trail
blockerWhy did we pick that approach? The answer was buried in a transcript on one of two vendors' platforms.
Someone else's sandbox
blockerCloud agent environments were slower than the PC under my desk and missing my tools.
I didn't need smarter agents. I needed a board, a process, and one place where the work and the decisions live.
Five days with Codex, and a fortress with no door
I had Codex usage to spare and wanted to throw something real at it. I was also trying out ChatGPT's new Dot feature, which orchestrates threads for you much like Claude Code Projects does. I didn't have access to Projects yet, so this was a chance to see how far that style of orchestration could go.
I knew a few days in that this was a swamp. But the tokens were there, and I was genuinely curious how far I could take it. And sure, the temptation to play the hero and rescue a beaten horse is always there.
At 9:31 PM on October 5, I told it to publish its last commit, "then stop all Software Factory development." My Claude usage had just reset, and it was time to call in my daily driver.
An honest code review, then an MVP in about fifteen minutes
I think we've made it too complicated with too many checks/constraints/validations/authorizations/etc. … Can you do a full review of this project and let me know your thoughts on whether it's over engineered, under engineered, really great, or awful?
"Over-engineered in the control plane, under-engineered in the product, and the two are directly related."
The code wasn't broken. It was "guarding a door nobody can walk through." It treated the agent as an adversary of the operator, when the operator was me, alone, on my own machines.
The clock
- The verdict arrives.
- The rebuild merges. About fifteen minutes turned days of churn into a simple core that works.
- My PC claims its first real job: Design on a game ticket.
- First ticket ships end to end, through Design, Implementation and Testing, about four hours after the rebuild started.
An agile board staffed by agents
Six columns: Backlog, Design, Implementation, Testing, Final Review, Done. Agents work the middle three. Each phase can run on Claude Code or Codex, and the agents run on my own PCs, which check in for work, reserve a job and report back.
A review canvas instead of a chat box
I review designs on a canvas built for marking up: strike through lines, circle what matters, drop sticky notes. Send back hands it all to the same Design agent, with full context.
Eight versions in about five hours
EKDM-16 replaces regard words like "Cool" and "Cordial" with a ring. Version 4 used color alone, and my sticky note said what was wrong. Two versions later the agent drew the idea from my sketch. I circled the leftover brown bands and wrote one line.
The team talks in writing
Every ticket has a Discussions thread where agents question each other, and tag me when they're stuck.
Anatomy of a thread: 40 minutes from problem to decision
Here's one from EKDM-7, the home rentals ticket. The implementer found a problem it shouldn't solve alone and brought it to me. I agreed, then pulled in the Design agent, which I run at a higher reasoning effort than the implementer on purpose, to work it through with me.
- Implementation flags it. Design v9 is built, but the simulation shows visitors settling in town dropping from 19 to 8. It stops and waits for me.
- I ask why before I weigh in.
- It traces the root cause home by home, and lays out four options with a recommendation.
- I agree, add how real players build, and tag @Design with two ideas of my own.
- Design has already updated the test plan, recommends one idea, and explains why to hold off on the other.
- I approve the direction. Design writes revision 10, which routes back to the same implementer.
One record, mine, across every agent
Artifacts and memory live in my own Google Cloud project, not with Anthropic, OpenAI or xAI. When an agent learns something worth keeping, it saves a short note, and the next agent on that project, Claude or Codex, starts from it. I can read, edit or delete any note.
The game, and the Factory improving itself
How I deliver agentic solutions now
I stopped managing agents like chat windows and went back to managing them like a team.
Process beats prompts
A board, clear phases and two human gates did more for quality than any clever instruction.
Write it down, somewhere you own
Discussions, docs, test reports and memory live in my cloud. Switching vendors no longer means losing the team's history.
Make escalation cheap
An agent tags me, I answer in one reply. A good manager is reachable without hovering.
Review like a reviewer
Striking a sentence and circling a mock beats three paragraphs of feedback in a prompt.
Run on your own iron
My PCs are faster than cloud sandboxes and already have my tools. A lease and a few limits keep them honest.
Right threat model, then MVP
Codex built a fortress against its own operator. Claude built the smallest thing that worked, and everything good grew from there.
The cards are moving. I'm back in the saddle.