Visual primers for product leadersยทModule 3 of 4

Spec-Driven Development

When AI agents write the code, the specification becomes the work. This module shows the loop, what a spec has to contain, where people stay in charge, and what a product leader's week looks like when prototypes take hours.

Colour keyDataModel / computePeopleTools & systemsRisk
01

Coding stopped being the bottleneck, so deciding what to build became the job

Agents can produce working code in minutes from a clear description. What they cannot do is know what you meant. The spec, once a formality, is now the artefact that determines the outcome, and verification against it is the new quality gate.

Before: people write the code Idea Spec Code Test Ship After: agents write the code Idea Spec Plan Code Verify Ship the long bar moves left Bar length is where calendar time goes, not effort. The total shrinks, and the human share moves into the spec and the verification. Skip the spec and you get 'vibe coding': fast, fun, and unverifiable.
The top row is the shape most teams grew up with. The bottom row is the shape when an agent implements: the spec and the verification absorb the time that coding used to take.
Teaching notesAsk the learner where their team's calendar goes today. Then ask: if coding took one hour, what would you need to have written down first? That list is the spec. Contrast with vibe coding, where you prompt, look, prompt again, and never write anything down.
Why a Head of Product cares
  • The PM's writing is now directly executable. Ambiguity that engineers used to resolve in conversation now becomes wrong code.
  • Speed shifts the risk from 'will it ship' to 'did we build the right thing'. Verification against the spec is the control.
  • Vibe coding is fine for throwaway prototypes and dangerous for anything that lives. Know which one you are doing.
Go deeper

Martin Fowler's site groups the current tools into a shared shape: a specification phase, a plan or design phase, a task breakdown, and implementation by an agent with checks. GitHub Spec Kit, Amazon's Kiro, and Tessl are the named examples; Claude Code's plan mode and project memory support the same shape without a dedicated framework. The common thread is that the spec is a living file in the repository, versioned with the code, rather than a document that goes stale after kickoff.

Basics refresher

A coding agent is an AI tool that can read and edit files, run commands, and iterate on its own work (Claude Code, Cursor, Copilot's agent mode). A repository (repo) is the version-controlled folder that holds a project's code and, increasingly, its specs. Vibe coding is prompting an agent to build something by feel, without a written spec or tests.

02

Spec-driven development is a seven-step loop with two ways back

People state intent and write the spec. The agent proposes a plan and a task list, which people approve. The agent implements. Verification checks the result against the spec's acceptance criteria, and failures go back either to the spec or to the code, never silently into production.

Intent the problem, in a breath 1 Spec what, why, criteria 2 Plan how: design, trade-offs 3 Tasks small, ordered, testable 4 Implement agent writes code + tests 5 Verify criteria met? tests pass? 6 Ship PR, review, deploy 7 people decide agent proposes, people approve machines check criteria were wrong or missing โ†’ fix the spec code is wrong โ†’ agent fixes
The three bands above the boxes are the division of labour. The two return arrows are the whole point: a failed check tells you whether the spec or the code was wrong, and each has a different fix.
Teaching notesWalk the loop with a small feature the learner knows. Stop at step 6 and ask: what would 'verified' mean here? If they cannot answer, the spec is missing acceptance criteria, which is the most common gap.
Why a Head of Product cares
  • Your leverage is steps 1, 2, and the approval in 3. Spend your time there.
  • Insist on step 6. An agent that says 'done' is reporting, not verifying; the check has to be independent of the thing that wrote the code.
  • The return arrows are your learning loop. Track how often a failure was a spec problem versus a code problem; it tells you where to invest.
Go deeper

In GitHub Spec Kit the phases are literally commands: specify, plan, tasks, implement. Kiro produces three files per feature: requirements (with acceptance criteria in a given/when/then style), design, and tasks. Claude Code's plan mode separates proposing from doing: the agent explores and writes a plan, the person approves, then the agent executes. In all of them the artefacts live in the repo next to the code and are reviewed like code. The Claude Code for PMs course frames the same idea for non-engineers: interview the person to build a requirements file before touching code (lesson 4.2, 'think of it like briefing a contractor').

Basics refresher

A pull request (PR) is a proposed change to a repository that others review before it is merged. Acceptance criteria are the specific, checkable conditions that mean a requirement is met. A task here is a unit of work small enough for an agent to finish and test in one go.

03

A good spec has eight parts, and an agent uses every one of them differently

A spec is not a longer prompt. Each section answers a distinct question, and the two sections PMs most often skip, non-goals and acceptance criteria, are the ones an agent depends on most: one fences the scope, the other becomes the tests.

Problem & why now the decision the reader should make anchors every trade-off the agent makes Users & scenarios who, doing what, in what situation personas the agent designs for Goals & success metrics measurable outcomes and guardrails what 'better' means; becomes the eval Non-goals what we are explicitly not doing the fence that stops scope creep Requirements functional and non-functional the checklist the plan must cover Acceptance criteria given / when / then, checkable turned directly into tests Constraints & dependencies systems, policies, deadlines limits the design space Open questions decisions still pending, with owners where the agent should ask, not guess left: the section ยท middle: what it says ยท right: what the agent does with it
Sections in the order a reader needs them. This merges standard PM specification practice with Lenny Rachitsky's and Carl Vellotti's PRD templates, trimmed to what an implementing agent can act on.
Teaching notesShow the learner a real PRD and tick off the eight sections. The usual finding is that non-goals and acceptance criteria are missing or vague. Rewrite one requirement as a given/when/then criterion together.
Why a Head of Product cares
  • Non-goals are the cheapest control you have over an agent. Without them, the agent will helpfully build things you did not ask for.
  • Acceptance criteria are where product and engineering meet. If you cannot write the check, the requirement is not ready.
  • Open questions with owners keep the agent from guessing. An agent that asks is worth more than one that assumes.
Go deeper

From Carl's template: problem and opportunity, high-level approach and alternatives considered, narrative, goals (measurable and not), non-goals with reasons, key features as the perimeter of the solution, and open questions. From Lenny's: description, problem, why now, success, audience, what it looks like, experiment plan, milestones. A fuller PM specification usually adds risks and mitigations, timeline, and communication plan; those stay in the PRD and are not needed by the agent. Non-functional requirements (latency, security, accessibility, cost) are where agents most often need explicit numbers, because the code will silently pick a default otherwise.

Basics refresher

A PRD is a product requirements document. Given / when / then is a way to write a testable criterion: given a state, when an action happens, then an outcome is observed. A non-functional requirement describes how well the system must work (speed, reliability) rather than what it does.

04

Three documents, three owners, one chain from why to steps

The PRD says why and what; the technical design says how; the task list says what to do first. Each layer is derived from the one above and reviewed by a different person. Agents are good at drafting downward and bad at inventing upward.

PRD why and what Technical design how: architecture, data, interfaces, trade-offs Task list ordered, small, each with its own test Owner Product lead writes, agent drafts sections Engineer or architect reviews; agent proposes Agent generates; engineer approves order Detail 1 to 3 pages as long as the trade-offs need dozens of one-line items
Detail grows downward; authority flows downward too. A task that cannot be traced to a requirement is scope creep, and a requirement with no task underneath it will not ship.
Teaching notesAsk who owns each layer on the learner's team today. In many teams the middle layer is missing (it lives in someone's head), which is exactly what breaks when an agent starts generating tasks.
Why a Head of Product cares
  • Own the top layer completely and review the middle layer's trade-offs. Delegate the bottom layer to the agent and check the order.
  • Traceability is a management tool: every task links to a criterion, so 'done' rolls up to the PRD automatically.
  • Missing middle layers show up as agents making architecture decisions by default. Make the engineer write or review the design before implementation.
Go deeper

Kiro's three files map exactly onto the pyramid (requirements, design, tasks). Spec Kit adds a project 'constitution' above the PRD: standing principles (testing policy, style, security rules) that every spec inherits. In Claude Code the same standing rules live in the project memory file (CLAUDE.md) and the design layer is what plan mode produces. Good task lists follow a pattern: each task names the files it touches, the test that proves it, and the criterion it serves; ordering puts data models and interfaces before behaviour before UI.

Basics refresher

An architecture is the set of components in a system and how they talk. An interface (API contract) is the agreed shape of a request and response between components. A data model is what gets stored and how it relates.

05

An agent only knows what you put in front of it, so curating that context is the new setup work

Project memory, company context, the spec, reusable skills, and tool access are the agent's entire world. Teams that maintain these files get consistent, on-brand output; teams that re-explain in every prompt get drift.

Coding agent Claude Code, Cursor... Project memory CLAUDE.md: stack, style, rules Company context product, personas, competitors The spec requirements + criteria Skills & commands reusable playbooks Tools via MCP repo, tickets, docs, browser Plan for approval Code + tests in a branch Pull request for review Reviewer approves plan, reviews PR steers what the agent knows = what you put in front of it; nothing else
Blue inputs are knowledge, green inputs are capabilities. The Claude Code for PMs course builds exactly this set: a CLAUDE.md with product context, personas, and writing style, plus company-context files and slash commands.
Teaching notesOpen the course's TaskFlow CLAUDE.md as an example: product context, three personas with quotes, writing style rules. Ask the learner what their equivalent file would contain, and who would keep it current.
Why a Head of Product cares
  • Context files are product assets. Personas, positioning, and tone written once are reused by every agent run and every teammate.
  • Stale context is worse than none: the agent will confidently apply last year's strategy. Assign an owner and a review cadence.
  • Tool access is a permissions decision, the same as for any agent in Module 1. Read before write; write behind review.
Go deeper

Project memory (CLAUDE.md or equivalent) holds the durable facts: stack, conventions, commands, what not to touch. Company context files hold product, personas, and competitive notes. Skills and slash commands package a repeatable workflow (write a PRD using these templates, review with three personas). Subagents are specialised agents with their own context (an engineer reviewer, an executive reviewer, a user researcher in the course). Together this is 'context engineering': deciding what goes into the context window, in what form, at what time. Raviv and Khan name context rot, the degradation as the window fills, as the reason to keep context lean and structured.

Basics refresher

A context window is the amount of text a model can see at once (Module 1). A slash command is a saved prompt invoked by name. Markdown is the plain-text format these files are written in; a heading is a line starting with #.

06

The tools are interchangeable; the stages are not

Every SDD tool serves one or more of the four stages. Some frameworks cover the whole loop; general-purpose agents cover it with a couple of conventions. Pick for the stage you are weakest at, and expect the list to change every six months.

Intent & spec Claude / ChatGPT for drafting PRD templates (Lenny, Carl) Claude Design for prototypes Plan & tasks Claude Code plan mode GitHub Spec Kit Kiro specs Implement Claude Code Cursor Copilot agent mode Verify & ship Tests in CI (GitHub Actions) Eval suites for AI features PR review, staged deploy the same person can run all four stages from one terminal; the stages, not the tools, are what matter
Tools that exist today, placed by stage. The prototype tool (Claude Design) sits under intent because a clickable prototype is a spec people can react to.
Teaching notesDo not teach tool features. Teach the learner to ask, for any new tool, which stage it serves and what artefact it leaves behind in the repo. Tools that leave nothing behind are for exploration only.
Why a Head of Product cares
  • Choose tools that write their artefacts into the repo. A spec inside a chat window is not a spec.
  • Prototype tools compress the intent stage: a spec can become an interactive HTML prototype within hours, with a one-click handoff to the coding agent.
  • Standardise the stages across teams, not the tools. Teams can differ on editors and still share PRD, design, and task conventions.
Go deeper

GitHub Spec Kit: an open-source CLI and templates with specify / plan / tasks / implement commands, working with many agents (Claude Code, Copilot, Cursor, Gemini CLI, Kiro and others). Kiro: Amazon's agentic IDE with specs as a first-class feature, generating requirements, design, and tasks files. Claude Code: plan mode for the design step, CLAUDE.md for memory, subagents and skills for reuse, and the ability to run tests and open pull requests. Claude Design: HTML-first prototyping grounded in a design system, with comments, in-place edits, and export to Claude Code as a zip with a README. Claude Cowork: the non-terminal path for non-technical teammates to run multi-step tasks on files.

Basics refresher

A CLI is a command-line interface, a tool you drive by typing commands. An IDE is an integrated development environment, the editor engineers work in. CI (continuous integration) runs tests automatically on every change; GitHub Actions is one such service.

07

Acceptance criteria are executable: write them so a test can check them

For ordinary features the criterion becomes a test the agent must make pass. For AI features it becomes an eval set with a pass-rate bar. Either way, 'done' is decided by a check the implementer did not write, and the PM wrote the criterion that check enforces.

Deterministic features Acceptance criterion given / when / then Test case written first Agent implements until tests pass Run tests in CI Pull request person reviews green red โ†’ fix AI features (probabilistic) Quality criterion 'answers cite a source' Eval set inputs + expected Agent tunes prompt, retrieval Score pass rate Ship gate score โ‰ฅ bar
Top lane: test-first development with an agent doing the implementing. Bottom lane: the same discipline for probabilistic features, where a single test cannot capture 'good' but a scored set can.
Teaching notesTake one requirement from the learner's backlog and write the criterion, then say out loud what the test would do. If it needs the word 'reasonable' or 'appropriate', it is not yet a criterion. For an AI feature, write three eval cases instead.
Why a Head of Product cares
  • This is the strongest lever a PM has over agent output: the criteria you write are the checks that will run.
  • Test-first with an agent is cheap. The agent writes the test from the criterion, then writes code until it passes; you review both.
  • For AI features, agree the pass-rate bar before building. It is the launch decision, made in advance.
Go deeper

Test-driven development with agents: the agent writes a failing test that encodes the criterion, implements, runs, and iterates; humans review the test as carefully as the code, because an agent can also weaken a test to make it pass. Contract tests pin interfaces between components. Property-based tests generate many inputs and check invariants. For AI features, evals use exact match, rubric grading by a stronger model with spot checks, or human review on samples, and they run in CI like tests so a prompt change that lowers the score blocks the merge. ByteByteGo's 'Good Code vs Bad Code' argues that readability and testability are what keep a codebase changeable, which matters more, not less, when an agent is the one changing it.

Basics refresher

A unit test checks one small piece of code in isolation; an integration test checks pieces working together. Red / green refers to failing and passing tests. A merge puts a reviewed change into the main codebase.

08

Three human gates and two automated gates keep speed from turning into risk

People decide the problem, the approach, and whether the code is right. Machines decide whether it passes and whether it is safe to roll out. Removing a human gate is a risk decision; removing an automated gate is a mistake.

Spec review right problem? Plan review right approach? Agent builds code + tests PR review right code? CI gate tests + evals? Staged deploy canary, then all Monitor did it work? each gate asks one question; a 'no' goes back one step, not to the start what we learn in production feeds the next spec orange gates are judgement calls a person owns; green gates are automated and should be boring
The pipeline shape most mature teams converge on. The three orange gates are where product and engineering judgement lives; everything green is ByteByteGo's CI/CD discipline applied to agent-written code.
Teaching notesAsk which gate the learner would remove to go faster, and what could go wrong. Spec review is the one teams most often skip and most often regret. Emphasise that PR review of agent code is a real skill: reviewers read for intent, not syntax.
Why a Head of Product cares
  • Decide the gates per risk tier. A marketing page can skip plan review; a payments change cannot.
  • PR review of agent-written code is where engineering time goes now. Budget for it and train for it.
  • Monitoring closes the loop back to the spec. The next spec should cite what the last one taught you.
Go deeper

From the CI/CD source: continuous delivery keeps every change deployable and promotes through environments (QA, staging, production) with checks at each; continuous deployment automates the last step. Staged rollouts (canaries, feature flags) limit blast radius. Code review of agent output focuses on: does it match the plan, are the tests real, did it touch files outside the task, did it introduce dependencies, are secrets and permissions handled. Review checklists are themselves good candidates for a skill file so the agent pre-reviews before a person does.

Basics refresher

A canary release sends a change to a small share of users first. A feature flag switches a feature on or off without redeploying. Staging is a production-like environment used for final checks.

09

A product leader can take a feature from idea to reviewable code in a day

Draft the PRD with company context loaded, generate a clickable prototype from it, gather comments on the live thing, hand it to a coding agent, and put a pull request in front of engineers. The artefacts are real and versioned; the time is mostly writing and reviewing.

09:00 Draft PRD with Claude + company context 10:30 Prototype Claude Design, from the spec 12:00 Comments team reacts on the live prototype 14:00 Hand off zip + README to Claude Code 16:00 Pull request engineers review real code next day Ship one day, one feature, one person driving the artefacts left behind: PRD, prototype, spec file, branch, PR what used to take two weeks of meetings now takes a day of writing and reviewing
Assembled from the Claude Design prototype-to-handoff workflow and the Claude Code for PMs course. Not every feature deserves this, but every PM should have done it once to recalibrate what a week is for.
Teaching notesRecommend the learner run this loop on something small and opinionated from their roadmap (a copy change, a settings placement, one new screen). Then debrief: what did they need to write down that they normally would not?
Why a Head of Product cares
  • Prototypes replace a lot of alignment meetings. People react to a clickable thing faster and more honestly than to a document.
  • The PM's output becomes code-adjacent: a spec file and a branch. Learn enough Git to open and read a pull request.
  • Engineers move from writing the first version to reviewing and hardening it. Their leverage goes up; their tolerance for vague specs goes down.
Go deeper

Prototyping practice: front-load context (screenshots, Figma exports, code snippets) so the prototype matches the real product; iterate with comments, in-place edits, and knobs rather than re-prompting; export as standalone HTML, PDF, editable slides, to Figma, or to Claude Code. The caution: the prototype is a starting point, and durable code means wiring to a backend, version control, and engineers extending it. The PM course adds the workflow around it: PRD with personas and company context, then three reviewer subagents (engineer, executive, user researcher) before stakeholders see it.

Basics refresher

Git is the version-control system under GitHub; a branch is a parallel line of changes; a pull request asks to merge a branch. A design system is the shared library of components and styles a product's screens are built from.

10

Five ways the loop breaks, and all five are visible if you look

Over-specifying wastes the agent's judgement; skipping the spec wastes yours. Drift makes the spec a lie; skipped verification makes 'done' a guess; and an agent that edits the spec to match its code has quietly taken over the product decision.

Intent Spec Plan Implement Verify Ship Over-specification spec longer than the code Spec drift code changed, spec did not No verification 'the agent said done' agent rewrites the spec to match what it built vibe-only: intent straight to code, nothing written down
Each red marker sits on the step it breaks. The last one, the return arrow from implement to spec, is the subtle one: it looks like diligence and is actually the agent grading its own homework.
Teaching notesAsk the learner which of the five they have already seen in a human team (all of them exist without AI). The point is that agents make each one faster and quieter, so the controls have to be explicit.
Why a Head of Product cares
  • Protect the spec's authorship. The agent may propose spec changes; a person accepts them, and the change is visible in the diff.
  • Set a length norm: a spec is one to three pages of decisions, not a novel. If the agent needs more, it needs a design doc, not a longer PRD.
  • Make verification independent: tests and evals written from the criteria, not by the same run that wrote the code.
Go deeper

Practical controls: keep the spec in the repo so drift shows up in pull-request diffs; require every task to cite a criterion; run tests and evals in CI where the agent cannot skip them; have the agent write a verification report that names each criterion and how it was checked; and give agents write access to spec files only through a reviewed PR. Over-specification usually means the PRD is doing the design doc's job; move the 'how' into the middle layer of the pyramid.

Basics refresher

A diff shows exactly what changed between two versions of a file. Scope creep is work growing beyond what was agreed. Tech debt is the future cost of shortcuts taken now.

11

One feature, traced from criteria to tasks

Real-time collaboration for a document editor: five acceptance criteria become five plan components and five ordered tasks, each with the test that proves it. The traceability is what lets a product leader read a task list and know whether the spec is covered.

Spec: acceptance criteria Plan: components Tasks: ordered, each with a test C1 2 to 10 users edit one document C2 see other users' cursors C3 90% of edits sync < 500 ms C4 conflict rate under 5% C5 works on web and mobile WebSocket service Presence + cursor broadcast CRDT for merging edits Sync metrics + alerting Mobile client hooks T1 channel per doc; test: 10 sockets join T2 broadcast cursor; test: 2 clients see it T3 integrate CRDT; test: concurrent edits T4 measure sync p90; test: < 500 ms T5 mobile hooks; test: iOS + Android every task names the criterion it serves and the test that proves it; C3 and C4 each need two components
Based on the draft feature spec in the Claude Code for PMs course. The spec's open questions (poor connections, dropped sockets, presence indicators, live history) would become either criteria or explicit non-goals before implementation starts.
Teaching notesHave the learner add the missing rows: turn one open question into a criterion (dropped connection reconnects within 5 seconds) and one into a non-goal (no live edit history in v1). Then ask which task an agent should do first and why (T1, because everything depends on the channel).
Why a Head of Product cares
  • Reading a task list against the criteria is a ten-minute review that catches most scope gaps before any code exists.
  • Numbers in criteria (500 ms, 5%, 10 users) are what make tasks testable. Vague criteria produce untestable tasks.
  • Open questions must be resolved into criteria or non-goals before implementation, or the agent will resolve them for you.
Go deeper

The original draft also lists a technical approach (WebSockets, operational transform or CRDT), success metrics that double as criteria, a four-phase timeline, and resourcing. In an SDD workflow the technical approach moves to the design layer where the OT-versus-CRDT trade-off is argued and decided, the timeline becomes task ordering, and each phase's exit criterion is a test. A reviewer should notice that C3 is a percentile (p90 under 500 ms) and needs a measurement task, which is why T4 exists.

Basics refresher

A WebSocket is a persistent two-way connection between a browser and a server, used for live updates. A CRDT is a data structure that lets several people edit the same thing at once and always converge to the same result. p90 means 90% of measurements were at or below the value.

Glossary

Spec-driven developmentWrite the spec, plan, and tasks first; the agent implements; verify against the spec.
Vibe codingPrompting an agent by feel with nothing written down; fine for throwaways.
Acceptance criterionA checkable condition that means a requirement is met.
Non-goalSomething explicitly out of scope, with the reason.
Plan modeAn agent mode that proposes an approach for approval before editing anything.
Project memoryA file (CLAUDE.md) of durable facts and rules the agent reads every time.
Context engineeringChoosing what the agent sees, in what form, at what time.
Skill / slash commandA saved, reusable workflow the agent runs by name.
SubagentA specialised agent with its own context, used for a sub-task or a review persona.
MCPModel Context Protocol: standard way to give an agent tools (repo, tickets, docs).
Pull requestA proposed change to a repository that people review before merging.
CIContinuous integration: automated build and tests on every change.
EvalA scored set of inputs and expected outputs; the test suite for AI features.
Spec driftCode and spec diverging because one changed without the other.
TraceabilityEvery task points at the criterion it serves.
Spec Kit / KiroFrameworks that formalise specify, plan, tasks, implement.

Questions to ask your team

  1. Where does the spec live, and is it in the same repository as the code?
  2. Which of the eight spec sections are present? Show me the non-goals and the acceptance criteria.
  3. Who approved the plan before the agent started implementing?
  4. For each task, which criterion does it serve and which test proves it?
  5. What ran to verify this: tests, evals, or the agent's own report?
  6. Which gates does this change pass through, and which did we skip on purpose?
  7. What is in our project memory file, who owns it, and when was it last reviewed?
  8. What tools can the agent write to without a person in the loop?
  9. What did the last shipped feature teach us that the next spec should cite?

Reading list

  • Claude Code for Product Managers (Carl Vellotti) โ€” ccforpms.com, free and interactive. Lessons on planning before building, project memory, PRD templates, and company-context files; the real-time collaboration spec in section 11 is adapted from it.
  • Understanding Spec-Driven Development: Kiro, spec-kit, and Tessl (martinfowler.com) โ€” martinfowler.com. The clearest neutral comparison of the frameworks.
  • GitHub Spec Kit โ€” github.com/github/spec-kit. Open-source toolkit: specify, plan, tasks, implement.
  • Kiro specs documentation โ€” kiro.dev/docs/specs. Requirements, design, and tasks files.
  • How to build AI product sense (Tal Raviv & Aman Khan, Lenny's Newsletter) โ€” Part 1 and Part 2. Why running the loop yourself builds judgement; context engineering and context rot.
  • ByteByteGo (Alex Xu) โ€” A Crash Course in CI/CD and Good Code vs. Bad Code. The automated gates in sections 7 and 8.