Agentic is a dial, not a switch, and most value sits in the middle
Andrew Ng's first move is to retire the argument about what 'counts' as an agent. Every workflow has a degree of autonomy, from a single prompt to a model that designs its own plan. The business question is which setting a given job needs.
Why a Head of Product cares
- Roadmaps should name the autonomy level, not just 'add an agent'. A fixed workflow and a planning agent have different costs, risks, and review needs.
- Vendors sell the far right; production runs in the middle. Ask any vendor demo which notch it is really at.
- You can move right incrementally. Start with a fixed workflow, then let the model choose among a few tools, then let it plan. Each step is measurable.
Go deeper
Ng's course describes three practical degrees: low (the developer codes the sequence of LLM calls and tool calls), medium (the LLM chooses which of a fixed set of tools or steps to run and in what order), and high (the LLM plans, may write code or new tools, and adapts when a step fails). Module 5 covers the high end under 'Patterns for highly autonomous agents'. The course's own summary: agentic workflows 'plan multi-step processes, execute them iteratively, and improve outputs through reflection and tool use'.
Basics refresher
A workflow is an ordered set of steps that turns an input into an output. An LLM call is one request to the model with a prompt. Hard-coded means the sequence is written by a developer and does not change at run time.
Letting a model iterate beats upgrading to a bigger model
Ng's headline evidence: wrapping an older model in a loop of drafting, testing, and revising lifted coding accuracy far more than the jump to the next-generation model did. The workflow, not the model, was the bigger lever.
Why a Head of Product cares
- Budget for workflow engineering before budgeting for frontier-model spend. The loop is cheaper and often more effective.
- Model upgrades still matter, but they compound with a good workflow rather than replacing one.
- The same logic applies to your own product: a smaller model in a well-designed loop can beat a bigger model called once.
Go deeper
The comparison comes from Ng's March 2024 letter in The Batch, using the HumanEval benchmark of coding problems. Other benefits the course lists: agentic workflows can run steps in parallel (fetch nine web pages at once and then synthesise), they are modular (swap a component without rewriting the system), and they can exceed human speed on repetitive multi-step jobs. The cost is more LLM calls per task, which is why Module 4 spends time on latency and cost.
Basics refresher
Zero-shot means one prompt, no examples, no retries. A benchmark is a fixed test set used to compare models; HumanEval is a set of programming problems with automatic checks. Accuracy here is the share of problems whose code passes the tests.
The first design skill is breaking a job into the steps a person would take
Before any code, write down how a competent person does the task: gather, outline, draft, critique, revise. Each step becomes an LLM call or a tool call, and each hand-off becomes a place you can inspect and measure.
Why a Head of Product cares
- Decomposition is a product and operations skill, not an engineering one. Whoever understands the work best should draw the steps.
- Explicit steps are what make the system measurable. A single mega-prompt cannot be debugged; a five-step workflow can.
- Parallel steps are where the speed gains hide. Look for independent lookups that today run one after another.
Go deeper
Ng's standard illustration is essay writing: research, outline, first draft, self-critique, revision, rather than asking for the essay in one shot. The course extends this to business processes: identify where a human would iterate (reflection), where they would look something up or act (tool use), and where the sequence is not knowable in advance (planning). Steps that are independent should run concurrently. The decomposition also decides your eval strategy: every step boundary is a place to add a component-level check (section 7).
Basics refresher
Concurrency or parallelism means running independent steps at the same time. A hand-off is the output of one step becoming the input of the next. A trace is the recorded sequence of calls and outputs for one run.
Four design patterns cover agentic AI, and two of them are safe to start with today
Reflection and tool use are mature and give reliable improvements. Planning and multi-agent collaboration are more powerful and harder to predict. Ng's advice is to master the first two, then earn the right to the second two.
Why a Head of Product cares
- Most production wins in 2026 are reflection plus tools inside a fixed workflow. That is a low-risk, high-return investment.
- Planning and multi-agent designs need heavier evaluation and monitoring. Budget them as research until evals say otherwise.
- The patterns are vocabulary for reviewing a proposal: which of the four does this design use, and why?
Go deeper
Ng's one-line definitions: Reflection, the LLM examines its own work to find ways to improve it. Tool use, the LLM is given tools such as web search, code execution, or any other function to gather information, take action, or process data. Planning, the LLM comes up with and executes a multi-step plan to achieve a goal. Multi-agent collaboration, more than one AI agent work together, splitting up tasks and discussing and debating ideas. The course builds each from first principles in Python before introducing frameworks, so the learner understands what the framework is hiding.
Basics refresher
A design pattern is a named, reusable way to structure a solution to a recurring problem. A framework is a library that implements patterns so you do not write the plumbing yourself. First principles means building the simplest version by hand to understand it.
Reflection is automated code review: the model critiques its own draft, then revises
Ask the model for a draft, then ask it to check that draft against explicit criteria, then ask it to rewrite. Two or three rounds usually beat one careful prompt. It works best when something outside the model, like a test or a database, supplies the feedback.
Why a Head of Product cares
- Reflection is the cheapest quality upgrade available: no new data, no new model, one extra prompt in the loop.
- The quality of the critique prompt is a product decision. It encodes what 'good' means for your users.
- External feedback turns reflection from opinion into fact. Wherever your output can be executed, rendered, or checked, wire that check in.
Go deeper
Ng's canonical flow for code: generate code for task X; prompt 'check the code carefully for correctness, style, and efficiency and give constructive criticism'; feed the critique back and ask for a rewrite; repeat. The Module 2 lesson 'Why not just direct generation?' shows the quality gap, and 'Evaluating the impact of reflection' measures it rather than assuming it. External feedback examples from the course: execute generated SQL against the database and return errors or empty results; run unit tests; render a chart and have the model inspect the image. Reflection can be one model talking to itself or two roles (author and reviewer), which is the bridge to multi-agent designs.
Basics refresher
A unit test runs a piece of code with known inputs and checks the outputs. SQL is the query language for databases; generating it from English is a common agent task. A rubric is a checklist of named criteria used to judge quality.
Tool use is how text becomes action: the model asks, your software does
The model never touches your database. It reads a description of each tool, decides one is needed, and emits a structured request; ordinary code executes it and hands back the result. MCP standardises that plug, and code execution turns any library into a tool.
Why a Head of Product cares
- Every tool is a permission decision. Inventory what the agent can read, what it can write, and what a person must approve.
- MCP is worth a standard: tools described once become reusable across agents and vendors, which reduces lock-in.
- Code execution replaces a backlog of custom integrations, but it needs a sandbox. Treat it like giving a contractor a terminal.
Go deeper
Mechanics from the course and Ng's Batch letter: tool descriptions (name, purpose, argument schema) are placed in the model's context; the model outputs a formatted call such as {tool: web-search, query: ...}; post-processing detects it, runs the function, and appends the result for the next model turn. With hundreds of tools, a retrieval step selects the relevant subset, the same idea as RAG. Code execution lets the model write Python that calls thousands of library functions, which the course recommends over hand-building a tool per edge case. MCP (Model Context Protocol) defines how tools are described and invoked so one server works with many agents; ByteByteGo's summary is that MCP handles tool access while agent-to-agent protocols handle agent communication.
Basics refresher
A function is a named piece of code with inputs and an output. An API lets one program call another. A sandbox is an isolated environment where generated code can run without reaching production systems. JSON is the structured text format tool calls are written in.
Evals come in four kinds, and a small set of twenty examples beats none
Score objectively when there is a right answer and subjectively with a model judging against yes/no criteria when there is not. Score the whole workflow to know if it works and individual components to know why. Twenty to fifty examples are enough to start.
Why a Head of Product cares
- An eval set is the product spec for an agentic feature. If you cannot write twenty examples with a pass criterion, the feature is not defined yet.
- Component-level evals are what make agentic systems manageable: they turn 'it's flaky' into 'step three fails 30% of the time'.
- LLM-as-judge is legitimate and cheap when the rubric is binary and spot-checked by people. Fuzzy scores are neither.
Go deeper
Course guidance in order: (1) define what success looks like, (2) build a mini-eval of 20 to 50 examples for fast signal, (3) pick objective checks (code compares to ground truth) where an answer is verifiable and subjective checks (an LLM judge with a structured rubric of discrete yes/no criteria such as 'has a clear title', 'correct axis labels') where it is not. End-to-end evals are expensive and tell you whether the system works; component-level evals isolate the underperforming step and direct effort precisely. Keep eval data separate from any examples used in prompts. Rerun the suite on every prompt, model, or tool change.
Basics refresher
Ground truth is the known correct answer for an example. An LLM judge is a model prompted to grade another model's output. Binary scoring means each criterion is either met or not, no partial credit.
Error analysis tells you which step to fix; without it, agent debugging is guesswork
Read a batch of traces, mark where each failed run first went wrong, and count by component. The tally tells you where to spend the next week. For the weak step there are four moves, in cost order: refine the prompt, switch the model, split the step, fine-tune.
Why a Head of Product cares
- Ask for the failure tally by component in every agent review. It is the single most informative artefact an AI team can show you.
- The four fixes have very different costs. Insist that prompt refinement and step splitting are exhausted before fine-tuning appears on a roadmap.
- Traces must be logged from day one. You cannot analyse errors you did not record.
Go deeper
Method from the course: build the workflow end to end quickly, run it on a small set, and manually inspect each run's trace, comparing each component's output to what a competent person would have produced. Tag the first failing component. Aggregate the tags to find the component that fails most often, then apply component-level evals to that step so improvement can be measured. The four remedies for an LLM component are refining the prompt, switching models, decomposing the step into smaller steps, and fine-tuning. Tool components fail differently (bad inputs, timeouts, missing permissions) and are fixed in code.
Basics refresher
A trace is the log of every call, input, and output in one run. Tagging is labelling each failed run with the component responsible. Fine-tuning is retraining a model on your own examples (Module 1, section 6).
Build fast, analyse honestly, repeat: the development process the course teaches
Do not design the perfect agent up front. Build a rough end-to-end version, look hard at where it fails, fix the biggest problem, and go around again. Optimise for speed of response before cost per token, because users leave slow agents and finance rarely leaves cheap ones.
Why a Head of Product cares
- Fund agent projects in short cycles with a working end-to-end version as the first milestone, not a design document.
- Ask for latency numbers per step. Parallelising independent steps is usually the largest, cheapest speed-up.
- Cost optimisation is a later-stage activity. Do it when the bill is material and evals can prove quality did not drop.
Go deeper
The course's two modes: build (create the end-to-end system quickly, accept rough edges) and analyse (examine outputs, read traces, run small evals on 10 to 20 samples). The cycle tightens over time as evals grow and components stabilise. On performance: latency typically matters more than cost because it degrades the user experience and bottlenecks the run; parallel execution of independent tasks (the nine-page fetch example) is the first optimisation. Cost levers, in order of quality risk: fewer unnecessary calls, caching, smaller models for easy components, quantised or distilled models (see Module 2, section 11).
Basics refresher
Latency is how long a user waits for a response. Throughput is how many requests a system handles per second. A milestone here is a working system, not a document.
Planning hands the model the sequence, which is powerful and the hardest to predict
Instead of a developer fixing the steps, the model writes a plan for the goal, executes it with tools, and replans when a step fails. Ng calls it a very powerful capability with less predictable results. The sweet spot is a model plan over a small set of well-tested, hard-coded steps.
Why a Head of Product cares
- Planning is the notch on the dial where review requirements change. Any planning agent with write tools needs approval gates and full tracing.
- Keep the plan's building blocks small and tested. The model chooses among steps you have already evaluated, so error analysis still works.
- Planning with code execution is the pragmatic version: the model composes existing libraries instead of your team building endless tools.
Go deeper
In the course's framing, planning is the LLM autonomously deciding the sequence of steps for a larger task, adapting when things do not go as expected. It is less mature than reflection and tool use and produces less predictable behaviour, so use it where the decomposition cannot be specified in advance. The course suggests combining LLM planning with hard-coded discrete steps so component-level error analysis remains possible. Planning with code execution: the model writes a short program that calls library functions (pandas for data, date libraries for calendars), which handles edge cases the team never enumerated; the program runs in a sandbox and errors are fed back for a rewrite (reflection again).
Basics refresher
Replanning is generating a new plan after a step fails. A library is a collection of ready-made functions (pandas is the standard one for tables in Python). An approval gate is a point where a person must confirm before the agent proceeds.
Multi-agent systems are an org chart for prompts, and they inherit org-chart problems
Give each agent one role and a focused prompt, then let a manager agent delegate, the same way you would staff a project. Focused agents perform better than one agent doing everything. Free-form group discussion between agents is the least predictable pattern, so start with pipelines and hierarchies.
Why a Head of Product cares
- Multi-agent designs are a way to manage complexity, not a way to add intelligence. Each agent is still a prompt plus tools plus evals.
- Prefer pipelines and hierarchies. They keep traces readable and error analysis possible; free discussion does not.
- Frameworks (AutoGen, CrewAI, LangGraph) are fine, but the course builds the pattern by hand first so the team knows what the framework hides.
Go deeper
Ng's three reasons to split work across agents: ablation studies (for example in the AutoGen paper) show multiple agents outperform one; an LLM prompted to focus on one thing at a time performs better; and roles are a useful abstraction for decomposing a problem, like a manager organising a team. His software example staffs a software engineer, a product manager, a designer, and a QA engineer, each with its own system prompt. In the course, specialists are exposed as callable tools so the manager can plan across them, producing non-linear execution. Communication patterns range from sequential handoffs to hierarchical delegation to group chat; the further toward free interaction, the more evaluation and tracing you need. Protocols: MCP for tools, A2A and similar for agent-to-agent messages.
Basics refresher
A role prompt tells the model who it is and what it is responsible for. An orchestrator or manager is the component that decides which agent runs next. A handoff is one agent's output becoming another's input.
Glossary
Questions to ask your team
- Which notch on the autonomy dial is this workflow at, and what changes if we move it one to the right?
- Show me the decomposition. Which steps would a good employee iterate on, and which look something up?
- Which of the four patterns does this design use, and why not the simpler one?
- What external feedback does reflection get: tests, execution results, or only the model's own opinion?
- List every tool the agent can call. Which ones write, and which of those need approval?
- How many examples are in the eval set, and which are objective versus judged?
- What is the failure tally by component from the last twenty traces?
- What is the end-to-end latency, and which steps could run in parallel?
- For any planning or multi-agent design: what stops it, who approves writes, and can we read the trace?
Reading list
- Agentic AI course (Andrew Ng, DeepLearning.AI, updated Sept 2026) โ the course this module follows: five modules, labs in Python. Take it on Coursera (graded assignments) or on the DeepLearning.AI platform (all 41 video lessons free to audit; labs and quizzes need the Pro plan).
- Andrew Ng on agentic design patterns (The Batch, DeepLearning.AI, 2024) โ Agentic workflows (the HumanEval comparison in section 2 and the four-pattern list); Reflection (section 5); Tool Use (section 6); Planning (section 10, including the Wikipedia fallback story); Multi-Agent Collaboration (section 11).
- Inside Andrew Ng's Agentic AI Course: 15 Learnings (10x Playbooks) โ 10xplaybooks.com. Learner notes used to confirm the Module 4 material on evals, error analysis, and latency.
- ByteByteGo (Alex Xu) โ MCP vs RAG vs AI Agents (one-diagram companion to section 6) and MCP vs A2A vs ACP (protocols for section 11).
- Foundational Large Language Models & Text Generation (Google, 2025) โ kaggle.com. The cost and latency trade-offs behind section 9.