What goes in
- Work items: stories in Shortcut and GitHub issues. Agents read the tracker through its MCP server, the same way I do.
- Questions: from Slack, through PRCHD's team lane.
- Production: Sentry, plus scheduled triage runs fired by PRCHD's triggers and automations.
- Research: store listings and competitor teardowns, before a product gets a spec.
- Design drafts: screens from Google Stitch and Figma, turned into a written UI spec before anyone builds them.
Why write a spec when an agent can just start?
Because a coding agent is the most remote colleague I've ever had. It wasn't on any of the calls, it starts every session fresh, and it only knows what's written down.
- The first commit is the spec. PRCHD's first commit was a 357-line product spec. Pehra Pro's was 14,736 lines across six documents. Sqwrl's was a 3,014-line PRD. None of them contained any code.
- Specs are living documents. About 158 of them, roughly 77,000 lines, across fourteen repositories. In PRCHD, 44% of all commits touch the specs.
- Plans are temporary. Decisions aren't. Every feature gets a plan with an "Open questions" section. When it's built, its decisions move into the specs and the plan is deleted. PRCHD has written 243 plans and deleted 227 of them.
-
Architecture is written once, as skills. How a backend is laid out (POST-only use-case APIs,
Resulttypes) and how a Flutter app is structured live in skills the agents load by name.
Which agent do I use?
All three. Claude Code is the main one, and Codex and Grok run beside it.
- PRCHD drives all three behind one event schema, so switching costs a dropdown.
- They all run inside PRCHD workspaces: one worktree, one branch, one session per feature. In one month, September to early October 2026, 314 Claude Code sessions ran in PRCHD workspaces and 4 ran in a plain terminal. 312 of 364 Codex threads ran in PRCHD worktrees. I built the tool, then I moved into it.
- MCP connects them to everything else: Shortcut for work items, Stitch for screens, Dart's server for Flutter, Playwright for a browser they can drive.
One feature, start to finish
- Start fresh. A new worktree and an empty context. The agent learns the project from the spec, not from yesterday's chat.
- Plan. The agent drafts the plan, the requirements, the checks and the open questions. I review it before any code exists, while a mistake still costs a sentence to fix.
- Build. The agent writes the code. Risky parts (security, payments, migrations) go in smaller pieces, and I check each one before it moves on.
- Review for intent. Does it do what the spec says? Would I put my name on it? I don't argue about naming. I read diffs on my phone, with notes pinned to lines.
- Verify. Tests, lint, the build, and the app actually running. Then I read the tests, because a test I don't understand hasn't proved anything to me.
- Fix the spec, then the code. When review turns up a problem, it's usually a gap in the plan.
- Merge. Decisions go into the specs, and the plan is deleted.
Two features, two worktrees, two agents. They never share a working directory, and they only meet when I merge.
What comes out
- About 2,945 commits in 2026 so far, across twenty repositories.
- Products: PRCHD, Snap Receipt, QuickPeek, 247 Track, Pehra Pro.
- Three open-source packages on pub.dev.
- Long reads, listed on the Thesis.
What would I change with a team?
It's tuned for one engineer directing agents. With a team, these change first.
- CI builds releases. It doesn't gate merges. The checks run locally, first by the agent and then by me. For one engineer, that's a trade I'll make. For a team, merge gates in CI come first.
- Pull requests are rare. Most work merges locally after review on my phone. With a team, the pull request comes back as the shared record of why a change happened.
Do the habits work with humans?
They came from working with people. Agents made them mandatory.
- Write the decision where the next person will find it. I've worked remotely since 2014, with teams in China and then Canada. A decision made on a call reaches the people on the call, and nobody else.
- Run the board, not just the code. For two and a half years I was Scrum Master for MiID Cloud's Android and iOS teams: the events, the board, and the backlog alongside the product owner.
- Review for intent, and own the release. At Vizzn I led a mobile team of three: I planned their stories, reviewed their merge requests and owned release builds from development to production.
- Make the call, then let the agents build it. At Velzosoft I make the architecture decisions and write them down as specs. Agents build from them, several features at a time, and nothing merges until I've reviewed it.
- Kill things with evidence, and delete them properly. When the data said Snap Receipt's users were tradespeople, the consumer features and their endpoints were removed from the product.
Defaults
What I choose unless there's a reason not to.
- Boring technology in the core, sharp technology at the edges. Nothing I can't operate at 2 a.m.
- One tool that does four jobs beats four tools that do one each. NATS in Pehra Pro: streams, request/reply, a key-value store and live fan-out.
- The model is configuration, not a dependency. Every AI call goes through a layer where swapping the model is a one-line change.
- Explain the mechanism. Don't ask for trust. If I can't say how a guarantee is enforced, I don't make it.
- Agents get the narrowest tools that do the job. Snap Receipt's connector only reads. FieldAudit Pro's only reads and adds, and a test fails the build if a tool imports an update or delete path.
- Cost is a design input. Cloud bills and token bills are outputs of the architecture. Design for them on day one.
- Observability before scale. I want to hear about a failure from the system before I hear about it from a customer.
Tools
| For | I use |
|---|---|
| Agents | Claude Code (main), Codex, Grok |
| Directing them | PRCHD, on iPhone and iPad |
| Editor | VS Code |
| MCP servers | Shortcut, Google Stitch, Dart, Playwright |
| Work tracking | Shortcut, GitHub issues |
| Design | Figma, Google Stitch |
| Errors | Sentry |
| Network | Tailscale |