Write It Down First
Spec-driven development with coding agents, and where I picked up the habit.
If you've used a coding agent for more than a few days, you've probably had this moment.
You spend an evening explaining your project: what it's for, what it should never do, why you picked the database you picked. The agent gets it, and the work is good.
The next morning you open a new session and it's all gone. The agent is polite and capable, and it has no idea who you are. Before long, it suggests the database you already ruled out.
That reminded me of something. I'd been working this way for years before agents existed. Just with people.
Where the habit comes from
At the end of 2014 I started working remotely, with a product team in China. From 2018 it was a team in Canada. Different time zones, no shared office, no hallway.
You learn one thing quickly in that setup. A decision made on a call reaches the people who were on the call, and nobody else. If it isn't written down somewhere, in a story, a review comment or a spec, then for most of the team it never happened.
So we wrote things down. At MiID Cloud I ran the scrum board for the Android and iOS teams. At Vizzn I led a team of three mobile developers, planning their stories and reviewing their merge requests. A good part of both jobs was putting decisions where the next person would find them.
A coding agent is the most remote colleague I've ever had. It wasn't on any of the calls. It starts every session fresh. It only knows what's written down.
So I do what remote teams taught me to do. I write it down first.
What changed when agents arrived
For most of my career, the limit was how much code a team could write. That isn't the limit anymore. Agents write code quickly, and a lot of it is good.
What they can't do is know what you meant. Who the customer is. What a client contract rules out. Which option you already tried and dropped. That still has to come from a person, and it has to be somewhere the agent can read it.
For me, that place is the spec: a few Markdown files in the repository, next to the code.
Isn't a good prompt enough?
For small jobs, yes. Renaming a folder of screenshots, checking one number in a database, sketching a quick chart. I prompt, look, adjust, and I'm done. If the job fits in a sentence, a sentence is the right tool. That's what people call vibe coding, and I do it without guilt.
The trouble starts when the job doesn't fit in a sentence. A short prompt for a big job doesn't remove the decisions in it. It leaves them for the agent to make. Which database. What happens when the network drops. What the product must never do. The agent will pick something, and it may pick something different tomorrow.
And a prompt lives in a chat. When the session ends, the chat goes with it. A spec stays in the repository. Every new session starts by reading it, and every feature I finish leaves it a little more accurate.
Where it has paid off
Every product I've built this year started with a spec. PRCHD's first commit was its product spec, with no code in it, and that spec is still what the code gets checked against. Pehra Pro's first commit was its specification and database schema.
The real payoff comes when you change your mind. Halfway through building FieldAudit Pro, the plan changed: from an app that worked offline with an iCloud backup, to one that works offline first and syncs across a team. That's a big change. I changed the spec first, and the code followed it. Every session after that worked from the new plan, because the new plan was the only one written down.
This website went the same way. Its first spec described a fairly standard portfolio: a big intro, project cards, case studies. It was fine. It also looked like everyone else's. So I rewrote it, and later rewrote it again. Each rewrite cost me a document. Building each version first would have cost me a website.
What goes in a spec for a coding agent?
Less than you might expect. I write down what the agent can't know, and leave out what it already does.
For the project as a whole, three things:
- Why it exists. Who it's for, the problem it solves, and what's out of scope.
- What it's built on, and why. The reason matters more than the name. An agent that knows why you chose PostgreSQL won't quietly add a second database.
- The order of work. A roadmap in small steps, each one small enough to review in a sitting.
Then each feature gets a short plan of its own: what we're building, the requirements it has to meet, and how we'll know it works. That last part is the easiest to skip and the most useful. It lets the agent check its own work before it hands it to me.
What I leave out: variable names, class names, which loop to use. The agent is good at those, and a spec full of them is a spec nobody reads.
It's also worth keeping two kinds of file apart. A rules file, CLAUDE.md for Claude Code or AGENTS.md for Codex, tells the agent how to behave in your project: the commands, the conventions, what not to touch. The spec says what you're building and what you've decided. I keep both, in separate files.
Who writes the spec?
We both do. I have the final say.
I describe what I want, then ask the agent to interview me. Its questions are often good ones: a tradeoff I hadn't named, a package that already does the job, an edge case I'd otherwise meet weeks later. I answer, I decide, and the agent writes it up.
Then I read every line. Wherever I left a gap, the agent filled it with a guess, and a guess in the spec turns into a fact in the code. This is the part that needs me.
When the spec needs changing later, I ask the agent to change it rather than editing it by hand. That way the plan, the requirements and the checks stay in step with each other.
How one feature gets built
One feature, one branch, and a clean start each time.
- Start fresh. A new branch in its own git worktree, and a new agent session with an empty context. It learns the project from the spec, not from yesterday's chat.
- Plan. The agent drafts the plan, the requirements and the checks. I review them before any code exists, while a mistake still costs a sentence to fix. Claude Code has plan mode for this. In Codex I start the session read-only.
- Build. The agent writes the code. For the risky parts, like security, payments or database migrations, it works in smaller pieces and I check each one before it moves on.
- Review. I read the diff for intent. Does it do what the spec says? Would I put my name on it? I don't argue about naming.
- Verify. Tests, lint, the build, and the app actually running. Then I read the tests, because a test I don't understand hasn't proved anything to me.
- Fix the spec, then the code. When review turns up a problem, it's often a gap in the plan. So the plan gets fixed first, then the code.
- Merge. Commit, merge, and tick it off the roadmap.
Because each feature has its own worktree, two agents can work on two features at the same time. They never share a working directory, and they only meet when I merge.
What happens between features?
After each merge I stop and look at the plan again. Is the next item on the roadmap still the right one? Did this feature teach me something the spec should say? If the spec needs to change, that change goes in on its own branch, so I can always tell which version of the spec produced which code.
This is also when the process itself gets better. If I've typed the same instruction three times, it becomes a skill: a short, written workflow the agent loads by name. Claude Code and Codex can both use the same skill, so it comes with me whichever one I'm working in.
And when a feature raises a bigger question, like a different database or a new direction, I write it into a backlog file rather than onto the roadmap. It isn't lost, and it doesn't sneak into the work in front of me.
How do you stop a coding agent losing context?
Mostly by not asking it to remember much.
A long session fills the context window, and a full context window is where mistakes creep in. So I clear the context between features, with /clear in Claude Code or a new session in Codex, and let the spec do the remembering. If something only exists in the chat, it will be lost sooner or later.
For a second opinion on a big change, I don't ask the session that wrote it. In Claude Code I start a few subagents, each with a fresh context, and ask them to read the change against the spec. They regularly catch things the first session missed, and none of their reading clutters the main session.
The same habit protects me too. Agents write code faster than I can read it, and anything I merge without understanding becomes my problem later. So I keep features small, commits small, and the breaks between them clean. If I can't review it properly, it isn't ready to merge.
Does spec-driven development work on an existing codebase?
Yes. In one way it's easier, because the spec already exists. It's just spread across the code, the commit history, a to-do file and somebody's head.
So the first step is to ask the agent to read the project and write the spec it should have had: why it exists, what it's built on, what's left to do. I review that like any other spec, and from then on the loop is the same.
The agent can read the code and the commits. The part in somebody's head still needs a person.
What happens when you switch coding agents?
Very little, and that's the point. I move between Claude Code and Codex, sometimes on the same project, and the spec doesn't care which one is reading it. It's Markdown in a repository, and every agent reads Markdown.
Models change every few months. The spec is the part of the project that doesn't have to change with them.
Remote teams taught me that a decision only counts once it's written down. Coding agents made that lesson matter more, not less. They're fast, they're capable, and they work from what's on the page.
So my job has shifted. Less typing. More deciding what we're building, writing it down clearly, and checking that what comes back matches.