AI is a set of nested rings, and today's excitement lives in the two innermost
When someone says 'AI' in a product meeting, they almost always mean a large language model or another generative model. Knowing which ring you are in tells you what is possible, what it costs, and who you need.
Why a Head of Product cares
- Vendors sell 'AI' at every ring; the price and the risk profile differ by an order of magnitude between a fraud model and an LLM agent.
- Hiring follows the ring: a data scientist for classical ML, an ML engineer for deep learning, an AI engineer for LLM products.
- The outer rings are mature and boring in a good way. The inner ring is changing every quarter, so roadmaps there need shorter bets.
Go deeper
Machine learning fits a function to data. Deep learning uses neural networks with many layers so the function can be very complex. Generative AI is deep learning trained to produce outputs in the same form as its inputs (text in, text out). LLMs are generative models trained on text, built on the transformer architecture from 2017; their defining trick is predicting the next token so well that useful behaviour emerges.
Diffusion models (images), speech models, and multimodal models sit in the GenAI ring beside LLMs. Since 2024 most frontier models are multimodal: one model reads text, images, and audio.
Basics refresher
A model is a file of numbers (parameters) plus the code to run them. Training sets those numbers using data. Inference is running the finished model to get an answer. Everything in this module is one of those two activities.
Machine learning flips the direction of programming
Instead of a person writing the rules, you show the computer examples and correct answers, and training discovers the rules. The output of training is the model, which then runs like ordinary software.
Why a Head of Product cares
- Your leverage moves from writing requirements to curating examples. The quality of the data set is the product decision.
- Rules are explainable and cheap to change; learned models are more capable but need evaluation, monitoring, and retraining budgets.
- Many 'AI features' should still be rules. Ask whether the problem has a pattern people cannot write down; if they can write it down, write it down.
Go deeper
Training is an optimisation: pick the parameters that make the model's outputs closest to the labelled answers, measured by a loss function. Module 2 shows the loop. The key consequence is that a model is only as good as the examples: if the labels are biased, late, or from a different population than production, the learned rules are wrong in ways nobody wrote down.
Basics refresher
A label is the correct answer attached to an example (this email is spam). Structured data is rows and columns; unstructured data is text, images, and audio. Classical ML thrives on structured data; deep learning made unstructured data tractable.
An LLM writes one token at a time by predicting what comes next
The model never 'knows' the answer up front. It turns your prompt into numbers, runs them through stacked attention layers, gets a probability for every possible next token, picks one, and repeats. Everything surprising about LLMs follows from that loop.
Why a Head of Product cares
- Cost and latency scale with tokens in and out. Pricing, prompt design, and context limits are all token budgets.
- Because output is sampled from probabilities, the same prompt can give different answers. Determinism is a setting (temperature), not a default.
- The model has no database lookup step. Facts are patterns in weights, so anything rare, recent, or private needs to be supplied in the prompt (see RAG).
Go deeper
Tokens are sub-word pieces; roughly 3 to 4 characters of English each. Embeddings map each token to a vector of a few thousand numbers. Each transformer layer runs self-attention (each position computes how much to weigh every other position, with multiple 'heads' looking for different relationships) followed by a feed-forward network; residual connections and layer normalisation keep training stable. Frontier models stack dozens of these layers. The final layer produces a score for every token in the vocabulary, turned into probabilities with softmax.
Sampling controls (from the Google prompt-engineering whitepaper): temperature 0 always picks the top token (greedy, deterministic); higher values flatten the distribution and increase variety. Top-K limits choice to the K most likely tokens; top-P keeps the smallest set whose probabilities sum to P. Use low temperature for extraction and classification, higher for brainstorming.
The context window is the maximum number of tokens the model can attend to at once (prompt plus output). Attention cost grows with the square of that length, which is why long contexts are slower and pricier.
Basics refresher
A vector is just a list of numbers; two vectors are 'close' when their numbers are similar. A probability distribution assigns a share of 100% to every option. Streaming means the interface shows tokens as they are produced rather than waiting for the full answer.
A chat assistant is a base model plus two rounds of human shaping
Pretraining makes a model that completes text. Fine-tuning on human-written examples teaches it to answer. Learning from human preferences teaches it to answer well and safely. Your product is the fourth stage, and it is the only one you control.
Why a Head of Product cares
- You are buying stages 1 to 3 as a commodity. Switching vendors is mostly a stage-4 cost, so avoid designs that hard-wire one model's quirks.
- Alignment is what makes the model refuse or hedge. When it refuses something legitimate, that is a stage-3 artefact you route around with prompting, not a bug you can fix.
- The behaviour you see in a demo is the lab's defaults. Your prompt, context, and tools change it more than the choice between the top three vendors.
Go deeper
Pretraining is unsupervised next-token prediction on a huge corpus; it produces a base model that continues text but does not naturally answer questions. Supervised fine-tuning (SFT) trains on curated prompt-response pairs so the model behaves like an assistant. RLHF (reinforcement learning from human feedback) collects human rankings of candidate answers, trains a reward model to imitate those preferences, then optimises the assistant against the reward model. Newer variants use AI feedback (RLAIF) or verifiable rewards (RLVR) for maths and code, which is how reasoning models are trained.
Parameter-efficient fine-tuning (LoRA and adapters) lets a product team do a small stage-2 style update on top of a vendor model, training a tiny fraction of the weights.
Basics refresher
Fine-tuning means continuing training on a smaller, targeted data set. Reinforcement learning trains by reward rather than by labelled answers: the model tries things and is nudged toward what scored well. A base model is what exists before any of that shaping.
There are four ways to make a model yours, and you climb them in order
Prompting fixes instructions. RAG fixes missing knowledge. Fine-tuning fixes style and behaviour. Agents add the ability to act. Each rung costs more and is harder to undo, so exhaust the cheaper rung first.
Why a Head of Product cares
- The ladder is a budgeting tool. Each rung roughly multiplies the engineering and evaluation cost of the previous one.
- Fine-tuning locks knowledge into a model version; RAG keeps it in a document store you can update tonight. Prefer the one you can change without retraining.
- Agents multiply both value and risk: every tool call is an action in the world that needs permissions, logging, and a way to undo.
Go deeper
Prompting covers system prompts, examples, and structured output formats. RAG retrieves relevant passages at request time and puts them in the prompt (next section). Fine-tuning adjusts weights on your examples; parameter-efficient methods make it affordable but it still needs a clean training set and a held-out evaluation set. Agents wrap the model in a loop with tools and memory (section 8). A fifth option, training your own model, is rarely justified outside labs.
Basics refresher
A system prompt is the standing instruction the product sends with every request, invisible to the end user. A tool in this context is any function the model can request the software to run (search, database query, send email).
RAG gives the model an open book instead of trying to teach it everything
Retrieval-augmented generation finds the few passages relevant to the question and pastes them into the prompt. The model then answers from what it was shown, which makes answers current, private, and citable.
Why a Head of Product cares
- RAG is the default first architecture for any 'ask questions about our stuff' feature. It is a data pipeline problem more than a model problem.
- Quality depends on chunking, retrieval, and document hygiene. A brilliant model with bad retrieval still hallucinates.
- Citations are the trust feature. Design the UI to show which passages were used so users can verify.
Go deeper
Passages are embedded with an embedding model into vectors; similarity is cosine distance. Production systems usually combine vector search with keyword search (hybrid retrieval) and add a reranker that re-scores the top candidates. Agentic RAG (ByteByteGo EP169) lets the model decide whether to retrieve, rewrite the query, or try another source when the first result is empty. Metadata filters (tenant, date, permission) are essential for enterprise deployments so a user never retrieves another user's documents.
Basics refresher
An index is a data structure built for fast lookup. A vector database stores embeddings and answers 'which stored vectors are closest to this one'. Chunking is splitting a long document into passages small enough to fit several into one prompt.
An agent is a model in a loop with tools, memory, and permission to act
A chatbot answers once. An agent keeps going: it plans, calls a tool, reads the result, and decides the next step until the goal is met. The loop is the product; the model is one component inside it.
Why a Head of Product cares
- Agents turn AI from an answer into an outcome, which is where the business value in the McKinsey data sits, and where the risk sits too.
- Every tool exposed to an agent is a permission decision. Read-only tools first; write tools behind approval; irreversible tools last or never.
- Evaluate agents on task completion, not on individual answers. The four dimensions below are a usable scorecard.
Go deeper
The loop is often called ReAct (reason, act, observe). Tool calling is a structured output: the model emits a function name and arguments, the software runs it and feeds the result back as a message. Memory can be short-term (the conversation), working notes (a scratch file), or long-term (a store the agent reads and writes). Multi-agent systems split work across specialised agents; MCP handles tool access while agent-to-agent protocols (A2A) handle their communication.
A four-part agent scorecard: Effectiveness (goal achievement rate, constraint adherence, failure-mode breakdown), Efficiency (steps to success, time to completion, tokens or cost per successful task), Reliability (variance across runs, regression tracking, recovery rate on retry), Trustworthiness (traceability of reasoning and tool use, safety catches before the user, whether a human reviewer would ship the output).
Basics refresher
An API is a way for one program to ask another to do something. A function call is one such request with named inputs. MCP is a protocol, meaning an agreed format, for describing tools and calling them, so the same tool works with many agents.
A prompt is a layered brief, and most quality problems are missing layers
Role and rules, supplied context, the task, and the required output shape are four separate things. When an answer is wrong, find the missing layer before blaming the model.
Why a Head of Product cares
- Prompts are product specs written in prose. They belong in version control with tests, not in a Slack thread.
- Structured output (JSON) is what lets the rest of your software trust the model; insist on it for anything automated.
- Few-shot examples are the cheapest quality lever you have. Curating five good examples beats most other interventions.
Go deeper
Techniques from the whitepaper: zero-shot (instruction only), one- and few-shot (include examples), system / role / contextual prompting, step-back (ask a general question first), chain of thought (ask for reasoning before the answer, which improves accuracy on multi-step problems), self-consistency (sample several reasoning paths and vote), ReAct (interleave reasoning with tool calls). Best practices: provide examples, design for simplicity, be specific about output, prefer instructions over constraints, control length, use variables, and re-test when the model version changes.
Basics refresher
JSON is a plain-text format for structured data (named fields and values) that software can parse reliably. Version control (Git) tracks every change to a text file so you can see who changed a prompt and roll it back.
Every AI failure mode has a place in the pipeline where it enters, and a layer that catches it
Hallucination, prompt injection, leakage, bias, and cost blow-ups are predictable, not mysterious. A guarded pipeline with input checks, output checks, and human review for high-stakes paths is the standard shape.
Why a Head of Product cares
- Governance is a product decision about which flows are high-stakes. Map each feature to a risk tier and give tiers a standard set of guards.
- Hallucination is not fixable, only bounded. Retrieval with citations, structured output, and evals reduce it; a disclaimer does not.
- Prompt injection means any text the model reads (web pages, emails, documents) can try to give it instructions. Treat model inputs like untrusted user input.
Go deeper
Input guards: classifiers for injection and abuse, allow-lists for tools, per-user rate limits and token budgets. Output guards: PII detection, policy classifiers, schema validation, groundedness checks that compare the answer to retrieved passages. Observability: log prompts, retrieved context, tool calls, and outputs so failures can be replayed. Evaluation (section 12) is how you know the rate of each failure before and after a change. Regulatory context (EU AI Act risk tiers, sector rules) usually maps neatly onto the same high-stakes classification.
Basics refresher
PII is personally identifiable information. Rate limiting caps how often something can be called. A classifier is a small model that sorts inputs into categories (safe / unsafe) and is much cheaper than the main model.
Almost everyone uses AI; few have moved the P&L, and the difference is workflow redesign
Usage is near-universal, agent experiments are common, and enterprise-level profit impact is still rare. The organisations that see it set growth goals, redesign workflows around AI, and put senior leaders on it.
Why a Head of Product cares
- Pilots are cheap and everyone has them; the scarce capability is redesigning a workflow end to end, which is a product leadership job.
- Set growth or innovation objectives alongside cost. In the survey, efficiency-only programmes were the ones least likely to report significant value.
- Pick two functions and scale agents there rather than spreading thin. No function in the survey had more than roughly one in ten organisations scaling agents.
Go deeper
Other findings worth quoting: about half of high performers intend to redesign workflows; 64% of all respondents say AI is enabling innovation; expectations for workforce size over the next year split roughly a third decrease, four in ten no change, one in ten increase; and a third of high performers commit more than 20% of their digital budgets to AI. The report defines high performers as organisations attributing more than 5% of EBIT to AI and reporting 'significant' value, about 6% of respondents.
Basics refresher
EBIT is earnings before interest and tax, a standard profit measure. A pilot is a limited trial; scaling means the capability is in routine use across a function.
Evals are the new acceptance tests, and writing them is the PM's job
You cannot unit-test a probabilistic system, but you can score it against a set of real inputs with known-good outputs. The eval set is the sharpest statement of what 'good' means, and it is what lets you change models without fear.
Why a Head of Product cares
- An eval set is the contract with engineering. It replaces 'make it good' with a number that goes up.
- Model upgrades become routine: run the evals, compare, ship. Without them every vendor release is a re-launch.
- Building product sense requires hands-on time in the loop. Raviv and Khan recommend using a coding agent for non-technical work precisely because it shows reasoning, tool calls, and the context window filling up.
Go deeper
Scoring methods: exact match for classification and extraction; rubric grading by a stronger model (LLM-as-judge) with spot checks; human review for a sample. Track pass rate per category so a regression in one intent is visible. Keep eval data separate from few-shot examples in the prompt (otherwise you are testing on the training set). For agents, score end-to-end task completion and the four scorecard dimensions from section 8 rather than individual messages. Concepts Raviv and Khan single out as product-sense builders: context rot (quality degrades as the window fills), memory, tool selection, and context engineering.
Basics refresher
An acceptance test checks that a feature does what the spec said. A regression is when something that used to work stops working after a change. A held-out set is data the system has never seen, used only for scoring.
Glossary
Questions to ask your team
- Which ring is this feature in: rules, classical ML, or an LLM? What would the rules version look like?
- What is the token budget per request, and what does that cost at ten times today's volume?
- Which rung of the ladder are we on, and what evidence says the cheaper rung is not enough?
- Where does the model's knowledge come from: its weights, retrieved documents, or the user? How fresh is each?
- What can the agent do without a person approving? Which of those actions are irreversible?
- What is our eval set, how many cases, and what was the score before and after the last change?
- Which flows are high-stakes, and what guards do they have that low-stakes flows do not?
- If our vendor doubled prices tomorrow, how long would it take to switch models?
- What in this feature would still be ours if a competitor used the same model?
Reading list
- Google Gen AI whitepapers (Kaggle 5-Day Gen AI Intensive, 2025) โ Foundational Large Language Models & Text Generation and Prompt Engineering (Lee Boonstra). The LLM paper is the reference for sections 3 to 5 (its trade-offs chapter is the best short read on cost and latency); the prompting paper covers the sampling settings, techniques, and best practices behind section 9.
- The State of AI in 2025 (McKinsey / QuantumBlack) โ mckinsey.com. Source of the adoption statistics and the high-performer traits in section 11.
- How to build AI product sense (Tal Raviv & Aman Khan, Lenny's Newsletter) โ Part 1 and Part 2. The argument and exercises behind section 12.
- Generative AI Bible (CB Insights) โ cbinsights.com. Market landscape and the early timeline.
- ByteByteGo (Alex Xu) โ visual companions to this module: How LLMs See the World (tokenisation, section 3); How Large Language Models Learn and How LLMs Learn from the Internet (section 4); How to Shrink a Language Model (quantisation and distillation, section 5); MCP vs RAG vs AI Agents (sections 7 and 8); RAG vs Agentic RAG (the step after section 7); MCP vs A2A vs ACP (how agents talk to tools and each other).