Visual primers for product leadersยทModule 4 of 4

Agentic AI

Follows Andrew Ng's DeepLearning.AI course 'Agentic AI' (Coursera, updated September 2026) at leadership altitude: the autonomy spectrum, why iteration beats a bigger model, the four design patterns, and the evaluation discipline that separates demos from products.

Colour keyDataModel / computePeopleTools & systemsRisk
01

Agentic is a dial, not a switch, and most value sits in the middle

Andrew Ng's first move is to retire the argument about what 'counts' as an agent. Every workflow has a degree of autonomy, from a single prompt to a model that designs its own plan. The business question is which setting a given job needs.

less autonomous more autonomous Single prompt one call, one answer 'summarise this' Fixed workflow developer fixes the steps draft โ†’ critique โ†’ revise Model picks tools LLM picks the next tool 'look this up, then reply' Model plans LLM designs the workflow 'research and report on X' 'Agentic' is an adjective and a dial, not a yes/no label Ng's framing: stop arguing about what counts as an agent; ask how much autonomy a workflow has most business value today sits in the middle two: reliable, measurable, still automated the far right is powerful and least predictable; earn your way there
Course Module 1, 'What is agentic AI?' and 'Degrees of autonomy'. Left of the dashed frame is what most teams already do; right of it is where predictability drops and evaluation becomes non-negotiable.
Teaching notesHave the learner place three internal workflows on the dial. Then ask what would change if each moved one notch right: what new failure could occur, and who would notice? That question sets up evals in section 7.
Why a Head of Product cares
  • Roadmaps should name the autonomy level, not just 'add an agent'. A fixed workflow and a planning agent have different costs, risks, and review needs.
  • Vendors sell the far right; production runs in the middle. Ask any vendor demo which notch it is really at.
  • You can move right incrementally. Start with a fixed workflow, then let the model choose among a few tools, then let it plan. Each step is measurable.
Go deeper

Ng's course describes three practical degrees: low (the developer codes the sequence of LLM calls and tool calls), medium (the LLM chooses which of a fixed set of tools or steps to run and in what order), and high (the LLM plans, may write code or new tools, and adapts when a step fails). Module 5 covers the high end under 'Patterns for highly autonomous agents'. The course's own summary: agentic workflows 'plan multi-step processes, execute them iteratively, and improve outputs through reflection and tool use'.

Basics refresher

A workflow is an ordered set of steps that turns an input into an output. An LLM call is one request to the model with a prompt. Hard-coded means the sequence is written by a developer and does not change at run time.

02

Letting a model iterate beats upgrading to a bigger model

Ng's headline evidence: wrapping an older model in a loop of drafting, testing, and revising lifted coding accuracy far more than the jump to the next-generation model did. The workflow, not the model, was the bigger lever.

GPT-3.5, single prompt (zero-shot) 48.1% GPT-4, single prompt (zero-shot) 67.0% GPT-3.5 in an iterative agentic workflow 95.1% share of HumanEval coding problems solved (Ng, The Batch, March 2024)
Course Module 1, 'Benefits of agentic AI'. Ng's own words: the improvement from GPT-3.5 to GPT-4 'is dwarfed by incorporating an iterative agent workflow'. Single measure, so one colour.
Teaching notesAsk the learner why the third bar is so much higher: the workflow lets the model check its work and try again, the way a person writes a draft and revises. Then ask what that means for the 'wait for the next model' strategy.
Why a Head of Product cares
  • Budget for workflow engineering before budgeting for frontier-model spend. The loop is cheaper and often more effective.
  • Model upgrades still matter, but they compound with a good workflow rather than replacing one.
  • The same logic applies to your own product: a smaller model in a well-designed loop can beat a bigger model called once.
Go deeper

The comparison comes from Ng's March 2024 letter in The Batch, using the HumanEval benchmark of coding problems. Other benefits the course lists: agentic workflows can run steps in parallel (fetch nine web pages at once and then synthesise), they are modular (swap a component without rewriting the system), and they can exceed human speed on repetitive multi-step jobs. The cost is more LLM calls per task, which is why Module 4 spends time on latency and cost.

Basics refresher

Zero-shot means one prompt, no examples, no retries. A benchmark is a fixed test set used to compare models; HumanEval is a set of programming problems with automatic checks. Accuracy here is the share of problems whose code passes the tests.

03

The first design skill is breaking a job into the steps a person would take

Before any code, write down how a competent person does the task: gather, outline, draft, critique, revise. Each step becomes an LLM call or a tool call, and each hand-off becomes a place you can inspect and measure.

Topic from the user Plan queries 3 to 5 searches Search + fetch page 1 Search + fetch page 2 Search + fetch page n Synthesise notes โ†’ outline Write report draft in parallel reflect: critique the draft, revise each box is a step a person would also take; each arrow is a place you can measure
Course Module 1, 'Task decomposition: identifying the steps in a workflow', shown on the course's own research-agent example (which learners run in 'Try the research agent'). Parallel branches are the speed win; the dashed loop is the quality win.
Teaching notesPick a real task from the learner's team (a weekly competitor summary, a customer reply). Whiteboard the steps a good employee takes. That whiteboard is the workflow; note where they would iterate and where they would look something up.
Why a Head of Product cares
  • Decomposition is a product and operations skill, not an engineering one. Whoever understands the work best should draw the steps.
  • Explicit steps are what make the system measurable. A single mega-prompt cannot be debugged; a five-step workflow can.
  • Parallel steps are where the speed gains hide. Look for independent lookups that today run one after another.
Go deeper

Ng's standard illustration is essay writing: research, outline, first draft, self-critique, revision, rather than asking for the essay in one shot. The course extends this to business processes: identify where a human would iterate (reflection), where they would look something up or act (tool use), and where the sequence is not knowable in advance (planning). Steps that are independent should run concurrently. The decomposition also decides your eval strategy: every step boundary is a place to add a component-level check (section 7).

Basics refresher

Concurrency or parallelism means running independent steps at the same time. A hand-off is the output of one step becoming the input of the next. A trace is the recorded sequence of calls and outputs for one run.

04

Four design patterns cover agentic AI, and two of them are safe to start with today

Reflection and tool use are mature and give reliable improvements. Planning and multi-agent collaboration are more powerful and harder to predict. Ng's advice is to master the first two, then earn the right to the second two.

predictability of results โ†’ power Reflection critique and revise Tool use search, act, run code Planning model chooses the steps Multi-agent specialists collaborate mature, reliable: start here powerful, less predictable the four patterns combine: a planner that calls tools and reflects, staffed by several agents
Course Module 1, 'Agentic design patterns', with Ng's stated view from The Batch that reflection and tool use are 'more reliable' than planning and multi-agent collaboration. Sections 5, 6, 10, and 11 take one pattern each.
Teaching notesKeep this slide up as the map for the rest of the module. The colour code carries over: violet patterns change how the model thinks, the green one connects it to the world. Ask which pattern each of the learner's use cases from section 1 would need first.
Why a Head of Product cares
  • Most production wins in 2026 are reflection plus tools inside a fixed workflow. That is a low-risk, high-return investment.
  • Planning and multi-agent designs need heavier evaluation and monitoring. Budget them as research until evals say otherwise.
  • The patterns are vocabulary for reviewing a proposal: which of the four does this design use, and why?
Go deeper

Ng's one-line definitions: Reflection, the LLM examines its own work to find ways to improve it. Tool use, the LLM is given tools such as web search, code execution, or any other function to gather information, take action, or process data. Planning, the LLM comes up with and executes a multi-step plan to achieve a goal. Multi-agent collaboration, more than one AI agent work together, splitting up tasks and discussing and debating ideas. The course builds each from first principles in Python before introducing frameworks, so the learner understands what the framework is hiding.

Basics refresher

A design pattern is a named, reusable way to structure a solution to a recurring problem. A framework is a library that implements patterns so you do not write the plumbing yourself. First principles means building the simplest version by hand to understand it.

05

Reflection is automated code review: the model critiques its own draft, then revises

Ask the model for a draft, then ask it to check that draft against explicit criteria, then ask it to rewrite. Two or three rounds usually beat one careful prompt. It works best when something outside the model, like a test or a database, supplies the feedback.

Task 'chart Q3 sales' Generate first draft Critique against named criteria Revise second draft Output repeat until the critique finds nothing new (2 to 3 rounds) External feedback: stronger than self-critique alone Run the code / SQL does it execute? Unit tests did it pass? Render the chart does it look right? feed errors back into the critique vague prompts ('improve this') do little; name the criteria: correctness, labels, tone, missing steps
Course Module 2, 'Reflection design pattern', including the chart-generation and SQL-generation labs. The top loop is self-critique; the bottom lane is 'Using external feedback', which the course shows to be the stronger version.
Teaching notesDemonstrate live if you can: ask a model for a chart title and axis labels, then ask it to critique its own output against a checklist, then revise. Point out that 'improve this' produces little; 'check for X, Y, Z' produces real fixes.
Why a Head of Product cares
  • Reflection is the cheapest quality upgrade available: no new data, no new model, one extra prompt in the loop.
  • The quality of the critique prompt is a product decision. It encodes what 'good' means for your users.
  • External feedback turns reflection from opinion into fact. Wherever your output can be executed, rendered, or checked, wire that check in.
Go deeper

Ng's canonical flow for code: generate code for task X; prompt 'check the code carefully for correctness, style, and efficiency and give constructive criticism'; feed the critique back and ask for a rewrite; repeat. The Module 2 lesson 'Why not just direct generation?' shows the quality gap, and 'Evaluating the impact of reflection' measures it rather than assuming it. External feedback examples from the course: execute generated SQL against the database and return errors or empty results; run unit tests; render a chart and have the model inspect the image. Reflection can be one model talking to itself or two roles (author and reviewer), which is the bridge to multi-agent designs.

Basics refresher

A unit test runs a piece of code with known inputs and checks the outputs. SQL is the query language for databases; generating it from English is a common agent task. A rubric is a checklist of named criteria used to judge quality.

06

Tool use is how text becomes action: the model asks, your software does

The model never touches your database. It reads a description of each tool, decides one is needed, and emits a structured request; ordinary code executes it and hands back the result. MCP standardises that plug, and code execution turns any library into a tool.

Tool descriptions name, what it does, arguments 'search_web(query)', 'get_orders(customer_id)' Model decides a tool is needed Function call {tool: "get_orders", id: 4821} in the prompt emits Your software runs it database, API, email, calendar result goes back into the context MCP one standard plug for tools describe a tool once; any MCP-aware agent can call it Code execution the universal tool Code execution: instead of building a tool for every edge case, let the model write a few lines that call a library (pandas, a date library) and run them Risk: anything the model can call, it can call wrongly. Read-only tools first, write tools behind approval, sandbox the code runner.
Course Module 3, 'Tool use': 'What are tools?', 'Creating a tool', 'Tool syntax', 'Code execution', and 'MCP'. The email-assistant lab is the course's worked example of a model reading, drafting, and sending through tools.
Teaching notesWalk one call end to end with a customer-service example: the model sees a get_orders tool, emits a call with the customer id, the system runs it, the model reads the orders and replies. Then ask: which tools in their business are read-only, and which ones change the world?
Why a Head of Product cares
  • Every tool is a permission decision. Inventory what the agent can read, what it can write, and what a person must approve.
  • MCP is worth a standard: tools described once become reusable across agents and vendors, which reduces lock-in.
  • Code execution replaces a backlog of custom integrations, but it needs a sandbox. Treat it like giving a contractor a terminal.
Go deeper

Mechanics from the course and Ng's Batch letter: tool descriptions (name, purpose, argument schema) are placed in the model's context; the model outputs a formatted call such as {tool: web-search, query: ...}; post-processing detects it, runs the function, and appends the result for the next model turn. With hundreds of tools, a retrieval step selects the relevant subset, the same idea as RAG. Code execution lets the model write Python that calls thousands of library functions, which the course recommends over hand-building a tool per edge case. MCP (Model Context Protocol) defines how tools are described and invoked so one server works with many agents; ByteByteGo's summary is that MCP handles tool access while agent-to-agent protocols handle agent communication.

Basics refresher

A function is a named piece of code with inputs and an output. An API lets one program call another. A sandbox is an isolated environment where generated code can run without reaching production systems. JSON is the structured text format tool calls are written in.

07

Evals come in four kinds, and a small set of twenty examples beats none

Score objectively when there is a right answer and subjectively with a model judging against yes/no criteria when there is not. Score the whole workflow to know if it works and individual components to know why. Twenty to fifty examples are enough to start.

objective subjective end-to-end component code checks ground truth did the final answer match? did the SQL return the right rows? did the tests pass? LLM judge with a binary rubric has a clear title: yes/no cites a source: yes/no tone matches brand: yes/no check one step in isolation did the query planner produce 3 to 5 relevant searches? cheap, fast, pinpoints the fault judge one step's output is this outline complete? is this critique specific? run on 20 to 50 examples start with a mini-eval of 20 to 50 examples; define success criteria before building
Course Module 1, 'Evaluating agentic AI', and Module 4, 'Evaluations' and 'Component-level evaluations'. The lab adds a component-level eval to the research workflow from section 3. Binary rubric items beat fuzzy 1-to-10 scores because they are reproducible.
Teaching notesConnect back to Module 1's eval loop: this is the same idea with the course's taxonomy. Have the learner write three binary rubric lines for a feature they own. Emphasise the course's rule of thumb: define success criteria first, build a mini-eval, then choose the eval type.
Why a Head of Product cares
  • An eval set is the product spec for an agentic feature. If you cannot write twenty examples with a pass criterion, the feature is not defined yet.
  • Component-level evals are what make agentic systems manageable: they turn 'it's flaky' into 'step three fails 30% of the time'.
  • LLM-as-judge is legitimate and cheap when the rubric is binary and spot-checked by people. Fuzzy scores are neither.
Go deeper

Course guidance in order: (1) define what success looks like, (2) build a mini-eval of 20 to 50 examples for fast signal, (3) pick objective checks (code compares to ground truth) where an answer is verifiable and subjective checks (an LLM judge with a structured rubric of discrete yes/no criteria such as 'has a clear title', 'correct axis labels') where it is not. End-to-end evals are expensive and tell you whether the system works; component-level evals isolate the underperforming step and direct effort precisely. Keep eval data separate from any examples used in prompts. Rerun the suite on every prompt, model, or tool change.

Basics refresher

Ground truth is the known correct answer for an example. An LLM judge is a model prompted to grade another model's output. Binary scoring means each criterion is either met or not, no partial credit.

08

Error analysis tells you which step to fix; without it, agent debugging is guesswork

Read a batch of traces, mark where each failed run first went wrong, and count by component. The tally tells you where to spend the next week. For the weak step there are four moves, in cost order: refine the prompt, switch the model, split the step, fine-tune.

Plan queries 2 of 20 runs wrong here Search + fetch 1 of 20 runs wrong here Synthesise 9 of 20 runs wrong here Write 3 of 20 runs wrong here Reflect 1 of 20 runs wrong here Read 20 traces; for each run, mark the first component whose output a person would reject fix the step with the most failures first, then re-run the evals Four ways to fix an LLM step Refine the prompt cheapest, try first Switch the model bigger, or a specialist Split the step two smaller steps Fine-tune last resort, needs data the course's core discipline: without this, debugging an agent is guesswork
Course Module 4, 'Error analysis and prioritizing next steps', 'More error analysis examples', and 'How to address problems you identify'. Numbers are illustrative; the shape is what matters.
Teaching notesThis is the section leaders most often skip and the one Ng says separates practitioners. Ask the learner how their team decides what to fix today. If the answer is 'whatever the last customer complained about', this is the alternative.
Why a Head of Product cares
  • Ask for the failure tally by component in every agent review. It is the single most informative artefact an AI team can show you.
  • The four fixes have very different costs. Insist that prompt refinement and step splitting are exhausted before fine-tuning appears on a roadmap.
  • Traces must be logged from day one. You cannot analyse errors you did not record.
Go deeper

Method from the course: build the workflow end to end quickly, run it on a small set, and manually inspect each run's trace, comparing each component's output to what a competent person would have produced. Tag the first failing component. Aggregate the tags to find the component that fails most often, then apply component-level evals to that step so improvement can be measured. The four remedies for an LLM component are refining the prompt, switching models, decomposing the step into smaller steps, and fine-tuning. Tool components fail differently (bad inputs, timeouts, missing permissions) and are fixed in code.

Basics refresher

A trace is the log of every call, input, and output in one run. Tagging is labelling each failed run with the component responsible. Fine-tuning is retraining a model on your own examples (Module 1, section 6).

09

Build fast, analyse honestly, repeat: the development process the course teaches

Do not design the perfect agent up front. Build a rough end-to-end version, look hard at where it fails, fix the biggest problem, and go around again. Optimise for speed of response before cost per token, because users leave slow agents and finance rarely leaves cheap ones.

Build end-to-end, quick and rough Analyse read traces, run mini-evals ship a first version in days, not months fix the weakest component, tighten the loop each cycle Optimise latency before cost slow agents lose users; token cost is usually a smaller number than it feels. Parallelise independent steps first. then: smaller models for easy steps, caching, fewer reflection rounds where evals say they add nothing
Course Module 4, 'Development process summary' and 'Latency, cost optimization'. The loop is the same shape as the eval loop in Module 1 and the SDD loop in Module 3; the course adds the discipline of analysing traces by hand.
Teaching notesContrast with a waterfall AI project: six months of architecture, then a demo that fails on real inputs. Ask the learner how quickly their team could have a rough end-to-end agent running on 20 real examples. If the answer is more than two weeks, the scope is too big.
Why a Head of Product cares
  • Fund agent projects in short cycles with a working end-to-end version as the first milestone, not a design document.
  • Ask for latency numbers per step. Parallelising independent steps is usually the largest, cheapest speed-up.
  • Cost optimisation is a later-stage activity. Do it when the bill is material and evals can prove quality did not drop.
Go deeper

The course's two modes: build (create the end-to-end system quickly, accept rough edges) and analyse (examine outputs, read traces, run small evals on 10 to 20 samples). The cycle tightens over time as evals grow and components stabilise. On performance: latency typically matters more than cost because it degrades the user experience and bottlenecks the run; parallel execution of independent tasks (the nine-page fetch example) is the first optimisation. Cost levers, in order of quality risk: fewer unnecessary calls, caching, smaller models for easy components, quantised or distilled models (see Module 2, section 11).

Basics refresher

Latency is how long a user waits for a response. Throughput is how many requests a system handles per second. A milestone here is a working system, not a document.

10

Planning hands the model the sequence, which is powerful and the hardest to predict

Instead of a developer fixing the steps, the model writes a plan for the goal, executes it with tools, and replans when a step fails. Ng calls it a very powerful capability with less predictable results. The sweet spot is a model plan over a small set of well-tested, hard-coded steps.

Goal 'refund if eligible' Model writes a plan steps it will take Step 1: look up order + policy Step 2: decide eligible? Step 3: act issue refund a step fails or returns something unexpected โ†’ replan Planning with code execution Goal 'top 5 products by margin' Model writes code pandas over the sales table Run it sandboxed Answer or fix and rerun libraries already handle the edge cases; the model composes them instead of you building a tool per case
Course Module 5, 'Planning workflows', 'Creating and executing LLM plans', and 'Planning with code execution', with the customer-service agent lab as the worked example. The bottom lane is planning where the 'tools' are library functions the model calls from code it writes.
Teaching notesTell Ng's demo story: a research agent's web-search tool failed mid-demo, and the agent switched to a Wikipedia tool he had forgotten he provided, and finished the task. That is the upside. Then ask what the downside would have been if the fallback tool were 'send email'.
Why a Head of Product cares
  • Planning is the notch on the dial where review requirements change. Any planning agent with write tools needs approval gates and full tracing.
  • Keep the plan's building blocks small and tested. The model chooses among steps you have already evaluated, so error analysis still works.
  • Planning with code execution is the pragmatic version: the model composes existing libraries instead of your team building endless tools.
Go deeper

In the course's framing, planning is the LLM autonomously deciding the sequence of steps for a larger task, adapting when things do not go as expected. It is less mature than reflection and tool use and produces less predictable behaviour, so use it where the decomposition cannot be specified in advance. The course suggests combining LLM planning with hard-coded discrete steps so component-level error analysis remains possible. Planning with code execution: the model writes a short program that calls library functions (pandas for data, date libraries for calendars), which handles edge cases the team never enumerated; the program runs in a sandbox and errors are fed back for a rewrite (reflection again).

Basics refresher

Replanning is generating a new plan after a step fails. A library is a collection of ready-made functions (pandas is the standard one for tables in Python). An approval gate is a point where a person must confirm before the agent proceeds.

11

Multi-agent systems are an org chart for prompts, and they inherit org-chart problems

Give each agent one role and a focused prompt, then let a manager agent delegate, the same way you would staff a project. Focused agents perform better than one agent doing everything. Free-form group discussion between agents is the least predictable pattern, so start with pipelines and hierarchies.

Manager agent plans, delegates, assembles Researcher gathers evidence Analyst finds the pattern Writer drafts the report Reviewer critiques, signs off each specialist is a tool the manager can call Communication patterns Sequential pipeline A โ†’ B โ†’ C, most predictable Hierarchical manager delegates, as above Group discussion agents debate freely, least predictable
Course Module 5, 'Multi-agentic workflows' and 'Communication patterns for multi-agent systems', with the market-research team lab. Ng's caution from The Batch: output quality is 'hard to predict, especially when allowing agents to interact freely'.
Teaching notesThe org-chart analogy lands with leaders: ask how they would staff a market-research project, then map each role to an agent with its own prompt and tools. Then ask which handoffs they would want to inspect. Those are the eval points.
Why a Head of Product cares
  • Multi-agent designs are a way to manage complexity, not a way to add intelligence. Each agent is still a prompt plus tools plus evals.
  • Prefer pipelines and hierarchies. They keep traces readable and error analysis possible; free discussion does not.
  • Frameworks (AutoGen, CrewAI, LangGraph) are fine, but the course builds the pattern by hand first so the team knows what the framework hides.
Go deeper

Ng's three reasons to split work across agents: ablation studies (for example in the AutoGen paper) show multiple agents outperform one; an LLM prompted to focus on one thing at a time performs better; and roles are a useful abstraction for decomposing a problem, like a manager organising a team. His software example staffs a software engineer, a product manager, a designer, and a QA engineer, each with its own system prompt. In the course, specialists are exposed as callable tools so the manager can plan across them, producing non-linear execution. Communication patterns range from sequential handoffs to hierarchical delegation to group chat; the further toward free interaction, the more evaluation and tracing you need. Protocols: MCP for tools, A2A and similar for agent-to-agent messages.

Basics refresher

A role prompt tells the model who it is and what it is responsible for. An orchestrator or manager is the component that decides which agent runs next. A handoff is one agent's output becoming another's input.

Glossary

Agentic workflowA process where an LLM completes some or all steps of a task, often iteratively, with tools.
Degree of autonomyHow much the model decides: from fixed steps, to choosing tools, to planning its own workflow.
Task decompositionBreaking a job into the steps a competent person would take.
ReflectionThe model critiques its own output against criteria and revises.
External feedbackResults from tests, code execution, or a database fed back into reflection.
Tool useThe model emits a structured request; software runs the function and returns the result.
Function callingThe formatted way a model asks for a tool to be run.
MCPModel Context Protocol, a standard for describing and calling tools.
Code executionThe model writes code that runs in a sandbox, turning libraries into tools.
PlanningThe model writes and executes a multi-step plan, replanning on failure.
Multi-agent collaborationSeveral role-specific agents split a task, coordinated by a manager or a pipeline.
EvalA scored set of examples; objective (code vs ground truth) or subjective (LLM judge with a rubric).
Component-level evalA check on one step's output in isolation.
Error analysisReading traces and tallying which component failed first, to prioritise fixes.
TraceThe recorded sequence of calls, inputs, and outputs in one run.
Build-analyse loopBuild end to end fast, inspect failures, fix the weakest step, repeat.

Questions to ask your team

  1. Which notch on the autonomy dial is this workflow at, and what changes if we move it one to the right?
  2. Show me the decomposition. Which steps would a good employee iterate on, and which look something up?
  3. Which of the four patterns does this design use, and why not the simpler one?
  4. What external feedback does reflection get: tests, execution results, or only the model's own opinion?
  5. List every tool the agent can call. Which ones write, and which of those need approval?
  6. How many examples are in the eval set, and which are objective versus judged?
  7. What is the failure tally by component from the last twenty traces?
  8. What is the end-to-end latency, and which steps could run in parallel?
  9. For any planning or multi-agent design: what stops it, who approves writes, and can we read the trace?

Reading list

  • Agentic AI course (Andrew Ng, DeepLearning.AI, updated Sept 2026) โ€” the course this module follows: five modules, labs in Python. Take it on Coursera (graded assignments) or on the DeepLearning.AI platform (all 41 video lessons free to audit; labs and quizzes need the Pro plan).
  • Andrew Ng on agentic design patterns (The Batch, DeepLearning.AI, 2024) โ€” Agentic workflows (the HumanEval comparison in section 2 and the four-pattern list); Reflection (section 5); Tool Use (section 6); Planning (section 10, including the Wikipedia fallback story); Multi-Agent Collaboration (section 11).
  • Inside Andrew Ng's Agentic AI Course: 15 Learnings (10x Playbooks) โ€” 10xplaybooks.com. Learner notes used to confirm the Module 4 material on evals, error analysis, and latency.
  • ByteByteGo (Alex Xu) โ€” MCP vs RAG vs AI Agents (one-diagram companion to section 6) and MCP vs A2A vs ACP (protocols for section 11).
  • Foundational Large Language Models & Text Generation (Google, 2025) โ€” kaggle.com. The cost and latency trade-offs behind section 9.