4 live 1 building on the bench: Pehra Pro, Snap Receipt last shipped

What happens between an idea and a merged change?

Less typing than you'd think. A lot more writing.

Here's my working setup, drawn as the system it is: what goes in, what processes it, and what comes out. Every box is real, and the numbers come from my own repositories and agent logs.

The diagram

What goes in

  • Work items: stories in Shortcut and GitHub issues. Agents read the tracker through its MCP server, the same way I do.
  • Questions: from Slack, through PRCHD's team lane.
  • Production: Sentry, plus scheduled triage runs fired by PRCHD's triggers and automations.
  • Research: store listings and competitor teardowns, before a product gets a spec.
  • Design drafts: screens from Google Stitch and Figma, turned into a written UI spec before anyone builds them.

Why write a spec when an agent can just start?

Because a coding agent is the most remote colleague I've ever had. It wasn't on any of the calls, it starts every session fresh, and it only knows what's written down.

  • The first commit is the spec. PRCHD's first commit was a 357-line product spec. Pehra Pro's was 14,736 lines across six documents. Sqwrl's was a 3,014-line PRD. None of them contained any code.
  • Specs are living documents. About 158 of them, roughly 77,000 lines, across fourteen repositories. In PRCHD, 44% of all commits touch the specs.
  • Plans are temporary. Decisions aren't. Every feature gets a plan with an "Open questions" section. When it's built, its decisions move into the specs and the plan is deleted. PRCHD has written 243 plans and deleted 227 of them.
  • Architecture is written once, as skills. How a backend is laid out (POST-only use-case APIs, Result types) and how a Flutter app is structured live in skills the agents load by name.

Which agent do I use?

All three. Claude Code is the main one, and Codex and Grok run beside it.

  • PRCHD drives all three behind one event schema, so switching costs a dropdown.
  • They all run inside PRCHD workspaces: one worktree, one branch, one session per feature. In one month, September to early October 2026, 314 Claude Code sessions ran in PRCHD workspaces and 4 ran in a plain terminal. 312 of 364 Codex threads ran in PRCHD worktrees. I built the tool, then I moved into it.
  • MCP connects them to everything else: Shortcut for work items, Stitch for screens, Dart's server for Flutter, Playwright for a browser they can drive.

One feature, start to finish

  1. Start fresh. A new worktree and an empty context. The agent learns the project from the spec, not from yesterday's chat.
  2. Plan. The agent drafts the plan, the requirements, the checks and the open questions. I review it before any code exists, while a mistake still costs a sentence to fix.
  3. Build. The agent writes the code. Risky parts (security, payments, migrations) go in smaller pieces, and I check each one before it moves on.
  4. Review for intent. Does it do what the spec says? Would I put my name on it? I don't argue about naming. I read diffs on my phone, with notes pinned to lines.
  5. Verify. Tests, lint, the build, and the app actually running. Then I read the tests, because a test I don't understand hasn't proved anything to me.
  6. Fix the spec, then the code. When review turns up a problem, it's usually a gap in the plan.
  7. Merge. Decisions go into the specs, and the plan is deleted.

Two features, two worktrees, two agents. They never share a working directory, and they only meet when I merge.

What comes out

What would I change with a team?

It's tuned for one engineer directing agents. With a team, these change first.

  • CI builds releases. It doesn't gate merges. The checks run locally, first by the agent and then by me. For one engineer, that's a trade I'll make. For a team, merge gates in CI come first.
  • Pull requests are rare. Most work merges locally after review on my phone. With a team, the pull request comes back as the shared record of why a change happened.

Do the habits work with humans?

They came from working with people. Agents made them mandatory.

  • Write the decision where the next person will find it. I've worked remotely since 2014, with teams in China and then Canada. A decision made on a call reaches the people on the call, and nobody else.
  • Run the board, not just the code. For two and a half years I was Scrum Master for MiID Cloud's Android and iOS teams: the events, the board, and the backlog alongside the product owner.
  • Review for intent, and own the release. At Vizzn I led a mobile team of three: I planned their stories, reviewed their merge requests and owned release builds from development to production.
  • Make the call, then let the agents build it. At Velzosoft I make the architecture decisions and write them down as specs. Agents build from them, several features at a time, and nothing merges until I've reviewed it.
  • Kill things with evidence, and delete them properly. When the data said Snap Receipt's users were tradespeople, the consumer features and their endpoints were removed from the product.

Defaults

What I choose unless there's a reason not to.

  1. Boring technology in the core, sharp technology at the edges. Nothing I can't operate at 2 a.m.
  2. One tool that does four jobs beats four tools that do one each. NATS in Pehra Pro: streams, request/reply, a key-value store and live fan-out.
  3. The model is configuration, not a dependency. Every AI call goes through a layer where swapping the model is a one-line change.
  4. Explain the mechanism. Don't ask for trust. If I can't say how a guarantee is enforced, I don't make it.
  5. Agents get the narrowest tools that do the job. Snap Receipt's connector only reads. FieldAudit Pro's only reads and adds, and a test fails the build if a tool imports an update or delete path.
  6. Cost is a design input. Cloud bills and token bills are outputs of the architecture. Design for them on day one.
  7. Observability before scale. I want to hear about a failure from the system before I hear about it from a customer.

Tools

ForI use
AgentsClaude Code (main), Codex, Grok
Directing themPRCHD, on iPhone and iPad
EditorVS Code
MCP serversShortcut, Google Stitch, Dart, Playwright
Work trackingShortcut, GitHub issues
DesignFigma, Google Stitch
ErrorsSentry
NetworkTailscale