Visual primers for product leadersยทModule 1 of 4

AI Foundations

What a model actually is, how a chat assistant gets built, and the four ways you can turn one into a product. Every idea is a picture first; open the panels when you want the mechanics or a refresher on the underlying tech.

Colour keyDataModel / computePeopleTools & systemsRisk
01

AI is a set of nested rings, and today's excitement lives in the two innermost

When someone says 'AI' in a product meeting, they almost always mean a large language model or another generative model. Knowing which ring you are in tells you what is possible, what it costs, and who you need.

Artificial intelligence any software that does something we'd call 'smart' MACHINE LEARNING learns rules from data instead of being told them DEEP LEARNING many-layered neural networks GENERATIVE AI produces new text, images, code, audio Large language models Examples at each ring AI: rules engines, route planning, chess ML: fraud scores, recommendations, demand forecasts Deep learning: face recognition, speech- to-text, self-driving perception GenAI: image generation, code assistants, synthetic voices LLMs: ChatGPT, Claude, Gemini, Llama Most 'AI' conversations in 2026 mean the two innermost rings.
Each inner ring is a special case of the one outside it. Classical ML still runs most fraud, ranking, and forecasting systems; LLMs are the newest and most expensive ring.
2017 Transformer paper 2019 GPT-2, GANs go mainstream 2020 GPT-3: prompting works 2022 ChatGPT launches 2023 GPT-4, open-weight Llama 2024 Multimodal + reasoning models 2025 Agents and MCP take off 2026 Agentic workflows at scale
The milestones that moved AI from research to product. The 2017 transformer architecture is the ancestor of every model on the right (timeline compiled from the CB Insights GenAI trend reports and public releases).
Teaching notesStart by asking the learner which ring their favourite product feature lives in. Point out that recommendations at Netflix are ML but not GenAI, and that a chatbot is an LLM product. Use the timeline to show how recent the innermost ring is: the transformer is younger than the iPhone.
Why a Head of Product cares
  • Vendors sell 'AI' at every ring; the price and the risk profile differ by an order of magnitude between a fraud model and an LLM agent.
  • Hiring follows the ring: a data scientist for classical ML, an ML engineer for deep learning, an AI engineer for LLM products.
  • The outer rings are mature and boring in a good way. The inner ring is changing every quarter, so roadmaps there need shorter bets.
Go deeper

Machine learning fits a function to data. Deep learning uses neural networks with many layers so the function can be very complex. Generative AI is deep learning trained to produce outputs in the same form as its inputs (text in, text out). LLMs are generative models trained on text, built on the transformer architecture from 2017; their defining trick is predicting the next token so well that useful behaviour emerges.

Diffusion models (images), speech models, and multimodal models sit in the GenAI ring beside LLMs. Since 2024 most frontier models are multimodal: one model reads text, images, and audio.

Basics refresher

A model is a file of numbers (parameters) plus the code to run them. Training sets those numbers using data. Inference is running the finished model to get an answer. Everything in this module is one of those two activities.

02

Machine learning flips the direction of programming

Instead of a person writing the rules, you show the computer examples and correct answers, and training discovers the rules. The output of training is the model, which then runs like ordinary software.

Traditional software Rules written by people Data Program runs Answers + Machine learning Data examples Answers labels Training finds the pattern Rules = the model + the model then runs like a program
Top row: the way software has always worked. Bottom row: the inputs and outputs swap places, and the discovered rules become the model you deploy.
Teaching notesUse a concrete pair: spam filtering by keyword rules versus spam filtering by 10,000 labelled emails. Ask what happens when spammers change tactics in each approach (rules go stale; the model can be retrained).
Why a Head of Product cares
  • Your leverage moves from writing requirements to curating examples. The quality of the data set is the product decision.
  • Rules are explainable and cheap to change; learned models are more capable but need evaluation, monitoring, and retraining budgets.
  • Many 'AI features' should still be rules. Ask whether the problem has a pattern people cannot write down; if they can write it down, write it down.
Go deeper

Training is an optimisation: pick the parameters that make the model's outputs closest to the labelled answers, measured by a loss function. Module 2 shows the loop. The key consequence is that a model is only as good as the examples: if the labels are biased, late, or from a different population than production, the learned rules are wrong in ways nobody wrote down.

Basics refresher

A label is the correct answer attached to an example (this email is spam). Structured data is rows and columns; unstructured data is text, images, and audio. Classical ML thrives on structured data; deep learning made unstructured data tractable.

03

An LLM writes one token at a time by predicting what comes next

The model never 'knows' the answer up front. It turns your prompt into numbers, runs them through stacked attention layers, gets a probability for every possible next token, picks one, and repeats. Everything surprising about LLMs follows from that loop.

Prompt your text Tokens word pieces Embeddings numbers Transformer N stacked layers attention: every token looks at every other Probabilities for next token Pick sample append the chosen token and repeat, one token at a time, until a stop token "Summarise this" [Summ][arise][ this] [0.12, -0.8, ...] the: 41% this: 22% a: 9% ...
Left to right is one step. The dashed return path is why responses stream word by word and why longer outputs cost more: each token is a full pass through the model.
Teaching notesHave the learner narrate the loop for a one-sentence prompt. Then explain that the model cannot revise earlier tokens, which is why 'think step by step' helps: it gives the model tokens to reason with before committing to an answer.
Why a Head of Product cares
  • Cost and latency scale with tokens in and out. Pricing, prompt design, and context limits are all token budgets.
  • Because output is sampled from probabilities, the same prompt can give different answers. Determinism is a setting (temperature), not a default.
  • The model has no database lookup step. Facts are patterns in weights, so anything rare, recent, or private needs to be supplied in the prompt (see RAG).
Go deeper

Tokens are sub-word pieces; roughly 3 to 4 characters of English each. Embeddings map each token to a vector of a few thousand numbers. Each transformer layer runs self-attention (each position computes how much to weigh every other position, with multiple 'heads' looking for different relationships) followed by a feed-forward network; residual connections and layer normalisation keep training stable. Frontier models stack dozens of these layers. The final layer produces a score for every token in the vocabulary, turned into probabilities with softmax.

Sampling controls (from the Google prompt-engineering whitepaper): temperature 0 always picks the top token (greedy, deterministic); higher values flatten the distribution and increase variety. Top-K limits choice to the K most likely tokens; top-P keeps the smallest set whose probabilities sum to P. Use low temperature for extraction and classification, higher for brainstorming.

The context window is the maximum number of tokens the model can attend to at once (prompt plus output). Attention cost grows with the square of that length, which is why long contexts are slower and pricier.

Basics refresher

A vector is just a list of numbers; two vectors are 'close' when their numbers are similar. A probability distribution assigns a share of 100% to every option. Streaming means the interface shows tokens as they are produced rather than waiting for the full answer.

04

A chat assistant is a base model plus two rounds of human shaping

Pretraining makes a model that completes text. Fine-tuning on human-written examples teaches it to answer. Learning from human preferences teaches it to answer well and safely. Your product is the fourth stage, and it is the only one you control.

1 ยท Pretrain Trillions of tokens web, books, code Base model predicts next token months, $10M to $1B+ 2 ยท Instruct Human-written Q&A demonstrations Instruct model follows instructions weeks 3 ยท Align Human rankings of candidate answers Aligned model helpful, safe, on-brand weeks, ongoing 4 ยท Serve Your prompt, tools, guardrails, UI Your product chat, agent, feature your team, continuous Reward model learns what humans preferred
Stages 1 to 3 happen inside the model lab. Stage 4 is where product teams live: the same aligned model becomes very different products depending on prompt, tools, and guardrails.
Teaching notesEmphasise the asymmetry: labs spend hundreds of millions on stages 1 to 3, and a product team can spend a week on stage 4 and still differentiate. Then ask: if the model is the same for every competitor, where does your moat come from? (Data, workflow, distribution, evals.)
Why a Head of Product cares
  • You are buying stages 1 to 3 as a commodity. Switching vendors is mostly a stage-4 cost, so avoid designs that hard-wire one model's quirks.
  • Alignment is what makes the model refuse or hedge. When it refuses something legitimate, that is a stage-3 artefact you route around with prompting, not a bug you can fix.
  • The behaviour you see in a demo is the lab's defaults. Your prompt, context, and tools change it more than the choice between the top three vendors.
Go deeper

Pretraining is unsupervised next-token prediction on a huge corpus; it produces a base model that continues text but does not naturally answer questions. Supervised fine-tuning (SFT) trains on curated prompt-response pairs so the model behaves like an assistant. RLHF (reinforcement learning from human feedback) collects human rankings of candidate answers, trains a reward model to imitate those preferences, then optimises the assistant against the reward model. Newer variants use AI feedback (RLAIF) or verifiable rewards (RLVR) for maths and code, which is how reasoning models are trained.

Parameter-efficient fine-tuning (LoRA and adapters) lets a product team do a small stage-2 style update on top of a vendor model, training a tiny fraction of the weights.

Basics refresher

Fine-tuning means continuing training on a smaller, targeted data set. Reinforcement learning trains by reward rather than by labelled answers: the model tries things and is nudged toward what scored well. A base model is what exists before any of that shaping.

05

Model choice is a quality, speed, and cost trade-off that you re-make every quarter

Small models are cheap and fast enough for narrow tasks; frontier models are for judgement and multi-step work. Most products use several, routed by task. The right answer changes as the frontier moves, so design for swapping.

cost per token and latency โ†’ quality โ†‘ Small / fast classification, routing, extraction Mid-size most production chat and summarisation Frontier reasoning, agents, hard judgement open-weight (run it yourself) distil or quantise frontier moves up every few months
Every point is a model tier, not a brand. Techniques like distillation and quantisation pull a capable model down and left; the frontier itself keeps rising, so last year's frontier is this year's mid-size.
Teaching notesAsk the learner to place three features of their product on this plane. Usually one belongs bottom-left (cheap classification) and one belongs top-right (a judgement call). Introduce 'routing': a small model decides which big model to call.
Why a Head of Product cares
  • A single-model architecture overpays. Route simple calls to small models and reserve frontier calls for the steps that need them.
  • Open-weight models trade vendor dependence for operational burden; the question is whether you have the team to run them.
  • Budget for re-evaluation: a quarterly model review is a real ritual in AI product teams because the menu changes that fast.
Go deeper

From the Google LLM whitepaper's trade-off section: quality vs latency/cost can often be improved with marginal quality loss by using a smaller model or a quantised one; the practical question is whether the model still performs your task, not whether it lost benchmark points. Latency vs cost is really latency vs throughput: batching more requests per GPU lowers cost per token but raises per-request wait, so bulk offline jobs and interactive chat want different settings.

Quantisation stores weights in 8- or 4-bit integers instead of 32-bit floats: smaller memory, faster arithmetic, usually tiny quality change. Distillation trains a small student model to imitate a large teacher, keeping much of the quality at a fraction of the cost. Output-preserving speedups include prefix caching (reuse computation for a repeated system prompt), speculative decoding (a small model drafts, the big one verifies), and batching.

Basics refresher

Latency is how long one request takes; throughput is how many requests a system handles per second. Open-weight means the trained parameters are downloadable, so you can host the model yourself; closed means you call the vendor's API. GPU chips do the massive parallel arithmetic that inference needs.

06

There are four ways to make a model yours, and you climb them in order

Prompting fixes instructions. RAG fixes missing knowledge. Fine-tuning fixes style and behaviour. Agents add the ability to act. Each rung costs more and is harder to undo, so exhaust the cheaper rung first.

Prompting change the instructions hours to first version RAG add your knowledge days to first version Fine-tuning change the behaviour weeks to first version Agents let it act with tools weeks+ to first version control & effort start at the left; move right only when the previous step cannot get you there
Most 'we need to fine-tune' requests are really RAG problems, and most RAG problems start as prompt problems. Climb only when evaluation says the current rung is the limit.
Teaching notesGive the learner three complaints and have them pick the rung: 'it doesn't know our refund policy' (RAG), 'it sounds too formal' (prompt first, fine-tune if that fails), 'it should file the ticket itself' (agent).
Why a Head of Product cares
  • The ladder is a budgeting tool. Each rung roughly multiplies the engineering and evaluation cost of the previous one.
  • Fine-tuning locks knowledge into a model version; RAG keeps it in a document store you can update tonight. Prefer the one you can change without retraining.
  • Agents multiply both value and risk: every tool call is an action in the world that needs permissions, logging, and a way to undo.
Go deeper

Prompting covers system prompts, examples, and structured output formats. RAG retrieves relevant passages at request time and puts them in the prompt (next section). Fine-tuning adjusts weights on your examples; parameter-efficient methods make it affordable but it still needs a clean training set and a held-out evaluation set. Agents wrap the model in a loop with tools and memory (section 8). A fifth option, training your own model, is rarely justified outside labs.

Basics refresher

A system prompt is the standing instruction the product sends with every request, invisible to the end user. A tool in this context is any function the model can request the software to run (search, database query, send email).

07

RAG gives the model an open book instead of trying to teach it everything

Retrieval-augmented generation finds the few passages relevant to the question and pastes them into the prompt. The model then answers from what it was shown, which makes answers current, private, and citable.

Offline: build the library (runs on a schedule) Your documents policies, tickets, wiki Chunk split into passages Embed passage โ†’ vector Vector store index of passages Online: answer one question (runs per request) Question Embed question โ†’ vector Retrieve top-k nearest Prompt question + passages LLM search the index answer with citations to the passages used
The top lane runs whenever documents change; the bottom lane runs on every question. The dashed arrow is the actual retrieval: nearest neighbours in vector space.
Teaching notesAsk what breaks if the top lane is stale (answers cite old policy) versus if retrieval returns the wrong passages (confident wrong answers). Both look like 'the AI is wrong' to users but have different owners.
Why a Head of Product cares
  • RAG is the default first architecture for any 'ask questions about our stuff' feature. It is a data pipeline problem more than a model problem.
  • Quality depends on chunking, retrieval, and document hygiene. A brilliant model with bad retrieval still hallucinates.
  • Citations are the trust feature. Design the UI to show which passages were used so users can verify.
Go deeper

Passages are embedded with an embedding model into vectors; similarity is cosine distance. Production systems usually combine vector search with keyword search (hybrid retrieval) and add a reranker that re-scores the top candidates. Agentic RAG (ByteByteGo EP169) lets the model decide whether to retrieve, rewrite the query, or try another source when the first result is empty. Metadata filters (tenant, date, permission) are essential for enterprise deployments so a user never retrieves another user's documents.

Basics refresher

An index is a data structure built for fast lookup. A vector database stores embeddings and answers 'which stored vectors are closest to this one'. Chunking is splitting a long document into passages small enough to fit several into one prompt.

08

An agent is a model in a loop with tools, memory, and permission to act

A chatbot answers once. An agent keeps going: it plans, calls a tool, reads the result, and decides the next step until the goal is met. The loop is the product; the model is one component inside it.

Goal 'reconcile these invoices' LLM reasons and decides Think / plan what next? Act call a tool Observe read the result Memory notes, past steps Tools APIs, files, DB chosen action result recall via MCP remember done? no โ†’ Result to person or ask for approval done? yes
MCP (Model Context Protocol) is the standard plug between the loop and the tools: describe a tool once and any MCP-aware agent can call it. Everything in orange is where a person enters or approves.
Teaching notesWalk one lap of the loop out loud with a real task (find duplicate invoices). Then ask where you would put a human approval gate. Point out that ByteByteGo's one-liner is useful: MCP is the plumbing, RAG is the memory, the agent is the manager.
Why a Head of Product cares
  • Agents turn AI from an answer into an outcome, which is where the business value in the McKinsey data sits, and where the risk sits too.
  • Every tool exposed to an agent is a permission decision. Read-only tools first; write tools behind approval; irreversible tools last or never.
  • Evaluate agents on task completion, not on individual answers. The four dimensions below are a usable scorecard.
Go deeper

The loop is often called ReAct (reason, act, observe). Tool calling is a structured output: the model emits a function name and arguments, the software runs it and feeds the result back as a message. Memory can be short-term (the conversation), working notes (a scratch file), or long-term (a store the agent reads and writes). Multi-agent systems split work across specialised agents; MCP handles tool access while agent-to-agent protocols (A2A) handle their communication.

A four-part agent scorecard: Effectiveness (goal achievement rate, constraint adherence, failure-mode breakdown), Efficiency (steps to success, time to completion, tokens or cost per successful task), Reliability (variance across runs, regression tracking, recovery rate on retry), Trustworthiness (traceability of reasoning and tool use, safety catches before the user, whether a human reviewer would ship the output).

Basics refresher

An API is a way for one program to ask another to do something. A function call is one such request with named inputs. MCP is a protocol, meaning an agreed format, for describing tools and calling them, so the same tool works with many agents.

09

A prompt is a layered brief, and most quality problems are missing layers

Role and rules, supplied context, the task, and the required output shape are four separate things. When an answer is wrong, find the missing layer before blaming the model.

What you send System: role and rules 'You are a claims assistant. Never guess.' Context: documents, examples policy excerpt, 3 worked examples Task: the instruction 'Classify this claim and explain why.' Output format JSON with fields: category, reason Model Output checked by code Knobs temperature 0 1+ deterministic โ† โ†’ varied top-p / top-k: how many candidates are eligible max tokens: length cap
The four layers map to the Google prompt-engineering whitepaper's system, contextual, and role prompting plus output structuring. The knobs control randomness and length, not quality.
Teaching notesShow a bad one-line prompt and rebuild it layer by layer with the learner. Then set temperature to 0 for the classification task and explain why a creative-writing feature would want the opposite.
Why a Head of Product cares
  • Prompts are product specs written in prose. They belong in version control with tests, not in a Slack thread.
  • Structured output (JSON) is what lets the rest of your software trust the model; insist on it for anything automated.
  • Few-shot examples are the cheapest quality lever you have. Curating five good examples beats most other interventions.
Go deeper

Techniques from the whitepaper: zero-shot (instruction only), one- and few-shot (include examples), system / role / contextual prompting, step-back (ask a general question first), chain of thought (ask for reasoning before the answer, which improves accuracy on multi-step problems), self-consistency (sample several reasoning paths and vote), ReAct (interleave reasoning with tool calls). Best practices: provide examples, design for simplicity, be specific about output, prefer instructions over constraints, control length, use variables, and re-test when the model version changes.

Basics refresher

JSON is a plain-text format for structured data (named fields and values) that software can parse reliably. Version control (Git) tracks every change to a text file so you can see who changed a prompt and roll it back.

10

Every AI failure mode has a place in the pipeline where it enters, and a layer that catches it

Hallucination, prompt injection, leakage, bias, and cost blow-ups are predictable, not mysterious. A guarded pipeline with input checks, output checks, and human review for high-stakes paths is the standard shape.

User input Input guard filter, classify Model + retrieval, tools Output guard PII, policy, format Human review for high-stakes only Prompt injection input tries to hijack Hallucination confident, wrong Data leakage private info out Bias skewed for some users Over-trust rubber-stamp review Runaway cost loops, long inputs Where each failure enters
Red callouts are failure modes; green boxes are the layers that catch them. Guards are ordinary code and small models, not magic. Governance is deciding which paths get which layers.
Teaching notesPick one failure per callout and ask who owns it: injection is security, hallucination is product plus evals, leakage is data governance, cost is engineering, bias is legal and product together. Then ask which paths in their product deserve a human in the loop.
Why a Head of Product cares
  • Governance is a product decision about which flows are high-stakes. Map each feature to a risk tier and give tiers a standard set of guards.
  • Hallucination is not fixable, only bounded. Retrieval with citations, structured output, and evals reduce it; a disclaimer does not.
  • Prompt injection means any text the model reads (web pages, emails, documents) can try to give it instructions. Treat model inputs like untrusted user input.
Go deeper

Input guards: classifiers for injection and abuse, allow-lists for tools, per-user rate limits and token budgets. Output guards: PII detection, policy classifiers, schema validation, groundedness checks that compare the answer to retrieved passages. Observability: log prompts, retrieved context, tool calls, and outputs so failures can be replayed. Evaluation (section 12) is how you know the rate of each failure before and after a change. Regulatory context (EU AI Act risk tiers, sector rules) usually maps neatly onto the same high-stakes classification.

Basics refresher

PII is personally identifiable information. Rate limiting caps how often something can be called. A classifier is a small model that sorts inputs into categories (safe / unsafe) and is much cheaper than the main model.

11

Almost everyone uses AI; few have moved the P&L, and the difference is workflow redesign

Usage is near-universal, agent experiments are common, and enterprise-level profit impact is still rare. The organisations that see it set growth goals, redesign workflows around AI, and put senior leaders on it.

Use AI regularly in at least one function 88% At least experimenting with AI agents 62% Report any EBIT impact from AI 39% Scaling an agentic system somewhere 23% 'High performers' (>5% of EBIT from AI) 6%
Share of respondents, McKinsey State of AI 2025 (November 2025, roughly 1,900 respondents). Single measure, so bars share one colour; read the gap between the first bar and the third.
Teaching notesLet the learner react to the gap between 88% and 39%. Then share the high-performer traits from the report: growth and innovation objectives rather than efficiency alone, workflow redesign, CEO-level ownership, and scaling agents in a few functions rather than piloting everywhere.
Why a Head of Product cares
  • Pilots are cheap and everyone has them; the scarce capability is redesigning a workflow end to end, which is a product leadership job.
  • Set growth or innovation objectives alongside cost. In the survey, efficiency-only programmes were the ones least likely to report significant value.
  • Pick two functions and scale agents there rather than spreading thin. No function in the survey had more than roughly one in ten organisations scaling agents.
Go deeper

Other findings worth quoting: about half of high performers intend to redesign workflows; 64% of all respondents say AI is enabling innovation; expectations for workforce size over the next year split roughly a third decrease, four in ten no change, one in ten increase; and a third of high performers commit more than 20% of their digital budgets to AI. The report defines high performers as organisations attributing more than 5% of EBIT to AI and reporting 'significant' value, about 6% of respondents.

Basics refresher

EBIT is earnings before interest and tax, a standard profit measure. A pilot is a limited trial; scaling means the capability is in routine use across a function.

12

Evals are the new acceptance tests, and writing them is the PM's job

You cannot unit-test a probabilistic system, but you can score it against a set of real inputs with known-good outputs. The eval set is the sharpest statement of what 'good' means, and it is what lets you change models without fear.

Define 'good' with real users Eval set inputs + expected Run prompt + model Score auto + human Inspect failures why did it miss? change the prompt, the data, or the model score clears the bar โ†’ ship Ship
The loop runs on every prompt change, model upgrade, and retrieval tweak. Tal Raviv and Aman Khan's argument: product sense in AI comes from running this loop yourself in a coding agent, where you can watch the model's reasoning and tool calls.
Teaching notesHave the learner write three eval cases for a feature they know: input, expected output, and how they would score it (exact match, rubric, human). Note that 20 hand-written cases beat zero automated ones.
Why a Head of Product cares
  • An eval set is the contract with engineering. It replaces 'make it good' with a number that goes up.
  • Model upgrades become routine: run the evals, compare, ship. Without them every vendor release is a re-launch.
  • Building product sense requires hands-on time in the loop. Raviv and Khan recommend using a coding agent for non-technical work precisely because it shows reasoning, tool calls, and the context window filling up.
Go deeper

Scoring methods: exact match for classification and extraction; rubric grading by a stronger model (LLM-as-judge) with spot checks; human review for a sample. Track pass rate per category so a regression in one intent is visible. Keep eval data separate from few-shot examples in the prompt (otherwise you are testing on the training set). For agents, score end-to-end task completion and the four scorecard dimensions from section 8 rather than individual messages. Concepts Raviv and Khan single out as product-sense builders: context rot (quality degrades as the window fills), memory, tool selection, and context engineering.

Basics refresher

An acceptance test checks that a feature does what the spec said. A regression is when something that used to work stops working after a change. A held-out set is data the system has never seen, used only for scoring.

Glossary

TokenA word piece, the unit models read and write and the unit you pay for.
Context windowMaximum tokens a model can consider at once, prompt plus answer.
EmbeddingA list of numbers representing meaning; similar things sit near each other.
TransformerThe 2017 architecture behind every modern LLM, built on attention.
AttentionThe mechanism that lets each token weigh every other token when computing its meaning.
InferenceRunning a trained model to get an output.
TemperatureSampling knob: 0 is deterministic, higher is more varied.
RAGRetrieval-augmented generation: fetch relevant passages, put them in the prompt.
Fine-tuningContinuing training on your own examples to change behaviour.
AgentA model in a loop with tools and memory that works toward a goal.
MCPModel Context Protocol, a standard way to describe and call tools.
HallucinationA confident, fluent, wrong output.
Prompt injectionText the model reads that tries to give it new instructions.
EvalA scored set of inputs and expected outputs used to measure quality.
GuardrailA check before or after the model that blocks unsafe or malformed content.
Open-weight modelA model whose parameters are published so you can run it yourself.

Questions to ask your team

  1. Which ring is this feature in: rules, classical ML, or an LLM? What would the rules version look like?
  2. What is the token budget per request, and what does that cost at ten times today's volume?
  3. Which rung of the ladder are we on, and what evidence says the cheaper rung is not enough?
  4. Where does the model's knowledge come from: its weights, retrieved documents, or the user? How fresh is each?
  5. What can the agent do without a person approving? Which of those actions are irreversible?
  6. What is our eval set, how many cases, and what was the score before and after the last change?
  7. Which flows are high-stakes, and what guards do they have that low-stakes flows do not?
  8. If our vendor doubled prices tomorrow, how long would it take to switch models?
  9. What in this feature would still be ours if a competitor used the same model?

Reading list