Coding stopped being the bottleneck, so deciding what to build became the job
Agents can produce working code in minutes from a clear description. What they cannot do is know what you meant. The spec, once a formality, is now the artefact that determines the outcome, and verification against it is the new quality gate.
Why a Head of Product cares
- The PM's writing is now directly executable. Ambiguity that engineers used to resolve in conversation now becomes wrong code.
- Speed shifts the risk from 'will it ship' to 'did we build the right thing'. Verification against the spec is the control.
- Vibe coding is fine for throwaway prototypes and dangerous for anything that lives. Know which one you are doing.
Go deeper
Martin Fowler's site groups the current tools into a shared shape: a specification phase, a plan or design phase, a task breakdown, and implementation by an agent with checks. GitHub Spec Kit, Amazon's Kiro, and Tessl are the named examples; Claude Code's plan mode and project memory support the same shape without a dedicated framework. The common thread is that the spec is a living file in the repository, versioned with the code, rather than a document that goes stale after kickoff.
Basics refresher
A coding agent is an AI tool that can read and edit files, run commands, and iterate on its own work (Claude Code, Cursor, Copilot's agent mode). A repository (repo) is the version-controlled folder that holds a project's code and, increasingly, its specs. Vibe coding is prompting an agent to build something by feel, without a written spec or tests.
Spec-driven development is a seven-step loop with two ways back
People state intent and write the spec. The agent proposes a plan and a task list, which people approve. The agent implements. Verification checks the result against the spec's acceptance criteria, and failures go back either to the spec or to the code, never silently into production.
Why a Head of Product cares
- Your leverage is steps 1, 2, and the approval in 3. Spend your time there.
- Insist on step 6. An agent that says 'done' is reporting, not verifying; the check has to be independent of the thing that wrote the code.
- The return arrows are your learning loop. Track how often a failure was a spec problem versus a code problem; it tells you where to invest.
Go deeper
In GitHub Spec Kit the phases are literally commands: specify, plan, tasks, implement. Kiro produces three files per feature: requirements (with acceptance criteria in a given/when/then style), design, and tasks. Claude Code's plan mode separates proposing from doing: the agent explores and writes a plan, the person approves, then the agent executes. In all of them the artefacts live in the repo next to the code and are reviewed like code. The Claude Code for PMs course frames the same idea for non-engineers: interview the person to build a requirements file before touching code (lesson 4.2, 'think of it like briefing a contractor').
Basics refresher
A pull request (PR) is a proposed change to a repository that others review before it is merged. Acceptance criteria are the specific, checkable conditions that mean a requirement is met. A task here is a unit of work small enough for an agent to finish and test in one go.
A good spec has eight parts, and an agent uses every one of them differently
A spec is not a longer prompt. Each section answers a distinct question, and the two sections PMs most often skip, non-goals and acceptance criteria, are the ones an agent depends on most: one fences the scope, the other becomes the tests.
Why a Head of Product cares
- Non-goals are the cheapest control you have over an agent. Without them, the agent will helpfully build things you did not ask for.
- Acceptance criteria are where product and engineering meet. If you cannot write the check, the requirement is not ready.
- Open questions with owners keep the agent from guessing. An agent that asks is worth more than one that assumes.
Go deeper
From Carl's template: problem and opportunity, high-level approach and alternatives considered, narrative, goals (measurable and not), non-goals with reasons, key features as the perimeter of the solution, and open questions. From Lenny's: description, problem, why now, success, audience, what it looks like, experiment plan, milestones. A fuller PM specification usually adds risks and mitigations, timeline, and communication plan; those stay in the PRD and are not needed by the agent. Non-functional requirements (latency, security, accessibility, cost) are where agents most often need explicit numbers, because the code will silently pick a default otherwise.
Basics refresher
A PRD is a product requirements document. Given / when / then is a way to write a testable criterion: given a state, when an action happens, then an outcome is observed. A non-functional requirement describes how well the system must work (speed, reliability) rather than what it does.
Three documents, three owners, one chain from why to steps
The PRD says why and what; the technical design says how; the task list says what to do first. Each layer is derived from the one above and reviewed by a different person. Agents are good at drafting downward and bad at inventing upward.
Why a Head of Product cares
- Own the top layer completely and review the middle layer's trade-offs. Delegate the bottom layer to the agent and check the order.
- Traceability is a management tool: every task links to a criterion, so 'done' rolls up to the PRD automatically.
- Missing middle layers show up as agents making architecture decisions by default. Make the engineer write or review the design before implementation.
Go deeper
Kiro's three files map exactly onto the pyramid (requirements, design, tasks). Spec Kit adds a project 'constitution' above the PRD: standing principles (testing policy, style, security rules) that every spec inherits. In Claude Code the same standing rules live in the project memory file (CLAUDE.md) and the design layer is what plan mode produces. Good task lists follow a pattern: each task names the files it touches, the test that proves it, and the criterion it serves; ordering puts data models and interfaces before behaviour before UI.
Basics refresher
An architecture is the set of components in a system and how they talk. An interface (API contract) is the agreed shape of a request and response between components. A data model is what gets stored and how it relates.
An agent only knows what you put in front of it, so curating that context is the new setup work
Project memory, company context, the spec, reusable skills, and tool access are the agent's entire world. Teams that maintain these files get consistent, on-brand output; teams that re-explain in every prompt get drift.
Why a Head of Product cares
- Context files are product assets. Personas, positioning, and tone written once are reused by every agent run and every teammate.
- Stale context is worse than none: the agent will confidently apply last year's strategy. Assign an owner and a review cadence.
- Tool access is a permissions decision, the same as for any agent in Module 1. Read before write; write behind review.
Go deeper
Project memory (CLAUDE.md or equivalent) holds the durable facts: stack, conventions, commands, what not to touch. Company context files hold product, personas, and competitive notes. Skills and slash commands package a repeatable workflow (write a PRD using these templates, review with three personas). Subagents are specialised agents with their own context (an engineer reviewer, an executive reviewer, a user researcher in the course). Together this is 'context engineering': deciding what goes into the context window, in what form, at what time. Raviv and Khan name context rot, the degradation as the window fills, as the reason to keep context lean and structured.
Basics refresher
A context window is the amount of text a model can see at once (Module 1). A slash command is a saved prompt invoked by name. Markdown is the plain-text format these files are written in; a heading is a line starting with #.
The tools are interchangeable; the stages are not
Every SDD tool serves one or more of the four stages. Some frameworks cover the whole loop; general-purpose agents cover it with a couple of conventions. Pick for the stage you are weakest at, and expect the list to change every six months.
Why a Head of Product cares
- Choose tools that write their artefacts into the repo. A spec inside a chat window is not a spec.
- Prototype tools compress the intent stage: a spec can become an interactive HTML prototype within hours, with a one-click handoff to the coding agent.
- Standardise the stages across teams, not the tools. Teams can differ on editors and still share PRD, design, and task conventions.
Go deeper
GitHub Spec Kit: an open-source CLI and templates with specify / plan / tasks / implement commands, working with many agents (Claude Code, Copilot, Cursor, Gemini CLI, Kiro and others). Kiro: Amazon's agentic IDE with specs as a first-class feature, generating requirements, design, and tasks files. Claude Code: plan mode for the design step, CLAUDE.md for memory, subagents and skills for reuse, and the ability to run tests and open pull requests. Claude Design: HTML-first prototyping grounded in a design system, with comments, in-place edits, and export to Claude Code as a zip with a README. Claude Cowork: the non-terminal path for non-technical teammates to run multi-step tasks on files.
Basics refresher
A CLI is a command-line interface, a tool you drive by typing commands. An IDE is an integrated development environment, the editor engineers work in. CI (continuous integration) runs tests automatically on every change; GitHub Actions is one such service.
Acceptance criteria are executable: write them so a test can check them
For ordinary features the criterion becomes a test the agent must make pass. For AI features it becomes an eval set with a pass-rate bar. Either way, 'done' is decided by a check the implementer did not write, and the PM wrote the criterion that check enforces.
Why a Head of Product cares
- This is the strongest lever a PM has over agent output: the criteria you write are the checks that will run.
- Test-first with an agent is cheap. The agent writes the test from the criterion, then writes code until it passes; you review both.
- For AI features, agree the pass-rate bar before building. It is the launch decision, made in advance.
Go deeper
Test-driven development with agents: the agent writes a failing test that encodes the criterion, implements, runs, and iterates; humans review the test as carefully as the code, because an agent can also weaken a test to make it pass. Contract tests pin interfaces between components. Property-based tests generate many inputs and check invariants. For AI features, evals use exact match, rubric grading by a stronger model with spot checks, or human review on samples, and they run in CI like tests so a prompt change that lowers the score blocks the merge. ByteByteGo's 'Good Code vs Bad Code' argues that readability and testability are what keep a codebase changeable, which matters more, not less, when an agent is the one changing it.
Basics refresher
A unit test checks one small piece of code in isolation; an integration test checks pieces working together. Red / green refers to failing and passing tests. A merge puts a reviewed change into the main codebase.
Three human gates and two automated gates keep speed from turning into risk
People decide the problem, the approach, and whether the code is right. Machines decide whether it passes and whether it is safe to roll out. Removing a human gate is a risk decision; removing an automated gate is a mistake.
Why a Head of Product cares
- Decide the gates per risk tier. A marketing page can skip plan review; a payments change cannot.
- PR review of agent-written code is where engineering time goes now. Budget for it and train for it.
- Monitoring closes the loop back to the spec. The next spec should cite what the last one taught you.
Go deeper
From the CI/CD source: continuous delivery keeps every change deployable and promotes through environments (QA, staging, production) with checks at each; continuous deployment automates the last step. Staged rollouts (canaries, feature flags) limit blast radius. Code review of agent output focuses on: does it match the plan, are the tests real, did it touch files outside the task, did it introduce dependencies, are secrets and permissions handled. Review checklists are themselves good candidates for a skill file so the agent pre-reviews before a person does.
Basics refresher
A canary release sends a change to a small share of users first. A feature flag switches a feature on or off without redeploying. Staging is a production-like environment used for final checks.
A product leader can take a feature from idea to reviewable code in a day
Draft the PRD with company context loaded, generate a clickable prototype from it, gather comments on the live thing, hand it to a coding agent, and put a pull request in front of engineers. The artefacts are real and versioned; the time is mostly writing and reviewing.
Why a Head of Product cares
- Prototypes replace a lot of alignment meetings. People react to a clickable thing faster and more honestly than to a document.
- The PM's output becomes code-adjacent: a spec file and a branch. Learn enough Git to open and read a pull request.
- Engineers move from writing the first version to reviewing and hardening it. Their leverage goes up; their tolerance for vague specs goes down.
Go deeper
Prototyping practice: front-load context (screenshots, Figma exports, code snippets) so the prototype matches the real product; iterate with comments, in-place edits, and knobs rather than re-prompting; export as standalone HTML, PDF, editable slides, to Figma, or to Claude Code. The caution: the prototype is a starting point, and durable code means wiring to a backend, version control, and engineers extending it. The PM course adds the workflow around it: PRD with personas and company context, then three reviewer subagents (engineer, executive, user researcher) before stakeholders see it.
Basics refresher
Git is the version-control system under GitHub; a branch is a parallel line of changes; a pull request asks to merge a branch. A design system is the shared library of components and styles a product's screens are built from.
Five ways the loop breaks, and all five are visible if you look
Over-specifying wastes the agent's judgement; skipping the spec wastes yours. Drift makes the spec a lie; skipped verification makes 'done' a guess; and an agent that edits the spec to match its code has quietly taken over the product decision.
Why a Head of Product cares
- Protect the spec's authorship. The agent may propose spec changes; a person accepts them, and the change is visible in the diff.
- Set a length norm: a spec is one to three pages of decisions, not a novel. If the agent needs more, it needs a design doc, not a longer PRD.
- Make verification independent: tests and evals written from the criteria, not by the same run that wrote the code.
Go deeper
Practical controls: keep the spec in the repo so drift shows up in pull-request diffs; require every task to cite a criterion; run tests and evals in CI where the agent cannot skip them; have the agent write a verification report that names each criterion and how it was checked; and give agents write access to spec files only through a reviewed PR. Over-specification usually means the PRD is doing the design doc's job; move the 'how' into the middle layer of the pyramid.
Basics refresher
A diff shows exactly what changed between two versions of a file. Scope creep is work growing beyond what was agreed. Tech debt is the future cost of shortcuts taken now.
One feature, traced from criteria to tasks
Real-time collaboration for a document editor: five acceptance criteria become five plan components and five ordered tasks, each with the test that proves it. The traceability is what lets a product leader read a task list and know whether the spec is covered.
Why a Head of Product cares
- Reading a task list against the criteria is a ten-minute review that catches most scope gaps before any code exists.
- Numbers in criteria (500 ms, 5%, 10 users) are what make tasks testable. Vague criteria produce untestable tasks.
- Open questions must be resolved into criteria or non-goals before implementation, or the agent will resolve them for you.
Go deeper
The original draft also lists a technical approach (WebSockets, operational transform or CRDT), success metrics that double as criteria, a four-phase timeline, and resourcing. In an SDD workflow the technical approach moves to the design layer where the OT-versus-CRDT trade-off is argued and decided, the timeline becomes task ordering, and each phase's exit criterion is a test. A reviewer should notice that C3 is a percentile (p90 under 500 ms) and needs a measurement task, which is why T4 exists.
Basics refresher
A WebSocket is a persistent two-way connection between a browser and a server, used for live updates. A CRDT is a data structure that lets several people edit the same thing at once and always converge to the same result. p90 means 90% of measurements were at or below the value.
Glossary
Questions to ask your team
- Where does the spec live, and is it in the same repository as the code?
- Which of the eight spec sections are present? Show me the non-goals and the acceptance criteria.
- Who approved the plan before the agent started implementing?
- For each task, which criterion does it serve and which test proves it?
- What ran to verify this: tests, evals, or the agent's own report?
- Which gates does this change pass through, and which did we skip on purpose?
- What is in our project memory file, who owns it, and when was it last reviewed?
- What tools can the agent write to without a person in the loop?
- What did the last shipped feature teach us that the next spec should cite?
Reading list
- Claude Code for Product Managers (Carl Vellotti) โ ccforpms.com, free and interactive. Lessons on planning before building, project memory, PRD templates, and company-context files; the real-time collaboration spec in section 11 is adapted from it.
- Understanding Spec-Driven Development: Kiro, spec-kit, and Tessl (martinfowler.com) โ martinfowler.com. The clearest neutral comparison of the frameworks.
- GitHub Spec Kit โ github.com/github/spec-kit. Open-source toolkit: specify, plan, tasks, implement.
- Kiro specs documentation โ kiro.dev/docs/specs. Requirements, design, and tasks files.
- How to build AI product sense (Tal Raviv & Aman Khan, Lenny's Newsletter) โ Part 1 and Part 2. Why running the loop yourself builds judgement; context engineering and context rot.
- ByteByteGo (Alex Xu) โ A Crash Course in CI/CD and Good Code vs. Bad Code. The automated gates in sections 7 and 8.