Visual primers for product leadersΒ·Module 2 of 4

ML Fundamentals

How a model learns from data, how you know whether it is any good, and how it gets from a notebook into a product. This is the ring underneath LLMs, and it still runs most of the AI in production. Sections 04 and 09 carry the diagnostic habits from Andrew Ng's Stanford CS229.

Colour keyDataModel / computePeopleTools & systemsRisk
01

Machine learning comes in three flavours, defined by what feedback the model gets

Supervised learning has the right answers; unsupervised learning has only the data; reinforcement learning has a reward. Almost every business use case is supervised, which means someone has to produce the labels.

Supervised Examples with labels Predictor new input β†’ label train email β†’ spam? claim β†’ fraud? photo β†’ cat? Most business ML. Needs labelled history. Unsupervised Unlabelled raw records Structure groups, outliers find customers β†’ segments logs β†’ anomalies No labels needed. Harder to judge 'correct'. Reinforcement Agent tries actions Environment responds action reward game moves ad bidding LLM alignment (RLHF) Learns by trial. Needs a reward signal.
Each panel shows the input on top, the trained artefact below, and the products that live there. The arrow is the training signal that differs between them.
Teaching notesAsk the learner to classify three features from their own product. Then ask where the labels for the supervised one come from and who pays for them: this is usually the hidden cost in an ML roadmap.
Why a Head of Product cares
  • The labelling budget is the real cost of supervised ML. Historical outcomes (did the customer churn?) are free labels; human judgements (is this review toxic?) are not.
  • Unsupervised results need a person to interpret them; plan for analyst time, not just model time.
  • Reinforcement learning is powerful but rare in products because it needs a safe place to try things. Its main mainstream use today is shaping LLMs.
Go deeper

Supervised tasks split into classification (discrete label) and regression (a number). Unsupervised methods include clustering (k-means), dimensionality reduction (PCA, UMAP), and anomaly detection. Self-supervised learning, where the label is derived from the data itself (predict the next word, fill the blank), is how LLMs are pretrained and blurs the first two categories. Semi-supervised and active learning reduce labelling cost by choosing which examples to label.

Generative vs discriminative is a second split inside supervised learning (CS229 lecture 5): discriminative models (logistic regression, trees, most neural networks) learn the boundary between classes directly, while generative models (Naive Bayes, Gaussian discriminant analysis) learn what each class looks like and ask which one a new input resembles. Generative models need less data and were the workhorse of spam filtering; discriminative models win once data is plentiful. This is a different sense of 'generative' from 'generative AI', which just means a model that produces content. Reinforcement learning is formalised as a Markov decision process: states, actions, a reward, and the rules for how actions change the world (lectures 16 to 20). The product decision is the reward function: it is the spec, and an agent optimises exactly what it says rather than what you meant. The unsupervised toolkit in the middle panel is k-means and Gaussian mixtures for grouping (lecture 13) and PCA for compressing many features into a few (lecture 15).
Basics refresher

A label is the known outcome for a training example. A feature is one input variable (age, amount, number of logins). A prediction is the model's output for a new example.

Tom Mitchell's definition, used in CS229 lecture 1: a program learns from experience E at task T if its performance P on T improves with E. Every ML project can be written as that triple, and if one of the three is missing the project is not ready.
02

An ML system is a loop, not a launch

Framing, data, training, evaluation, deployment, and monitoring feed each other. The model you ship is a snapshot; the loop is the asset, and the team that owns the loop owns the outcome.

Frame the problem what decision, what metric 1 Collect data logs, history, sources 2 Prepare & label clean, join, annotate 3 Train fit candidate models 4 Evaluate held-out test, business cost 5 Deploy serve behind an API 6 Monitor drift, latency, outcomes 7 Retrain new data, new labels 8 Most projects loop here several times before deploy
Step 1 is the product manager's step and the one most often skipped. Steps 2 and 3 take most of the calendar. Steps 7 and 8 are why ML has an operating cost after launch.
Teaching notesAsk the learner to guess the share of time per step; the usual answer surprises people (data preparation is often half or more). Then ask what happens if monitoring is skipped: the model silently degrades and nobody knows until a customer complains.
Why a Head of Product cares
  • Budget for the loop, not the model: data pipelines, labelling, evaluation, and monitoring are recurring costs.
  • Framing decides everything downstream. 'Predict churn' is not a frame; 'flag accounts to call this week, at most 200, ranked by expected saved revenue' is.
  • Retraining cadence is a product decision with a cost. Daily for fraud, quarterly for demand forecasting, when-evidence-says for most things.
Go deeper

Framing includes choosing the target variable (what exactly is being predicted), the unit of prediction, the decision the prediction feeds, and the metric that reflects business cost. Data work includes joining sources, handling missing values, deduplication, and defining the point in time at which features are known (to avoid leakage). Training compares several model families with cross-validation. Evaluation is done on a held-out test set and then in a shadow or A/B deployment. Monitoring tracks input drift, prediction drift, and, where available, realised outcomes.

Basics refresher

A pipeline is a chain of automated steps that turn raw data into a trained model or a prediction. Deployment means the model runs on a server that other software can call. An API endpoint is the address that software calls.

03

Training is rolling downhill on an error landscape, one small step at a time

A model starts with random weights. Training measures how wrong it is on the examples (the loss), then nudges every weight a little in the direction that reduces the error, and repeats millions of times. Stop too late and it memorises the examples instead of learning the pattern.

Gradient descent minimum: best weights each step: measure error, nudge every weight slightly downhill height = loss (how wrong the model is) Loss over training training steps β†’ loss training set validation set overfitting begins
Left: the landscape is the loss as a function of the weights; each step is gradient descent. Right: the gap between the training curve and the validation curve is overfitting, and the dashed line is where you should have stopped.
Teaching notesThe ball metaphor is enough for most leaders; the extra idea to land is the right-hand chart. A model that scores perfectly on its own training data has learned nothing you can trust. Ask how they would detect this in a vendor's demo (ask for results on data the vendor never saw).
Why a Head of Product cares
  • 'It scored 99% in testing' means nothing until you know which data it was tested on. Ask for held-out results.
  • Training cost is compute time multiplied by iterations; for deep learning that is GPU hours, and for LLMs it is the reason only labs pretrain.
  • Overfitting is the ML version of a demo that works only on the demo data. The defence is a validation set the team never tunes against.
Go deeper

The loss function turns predictions and labels into one number (cross-entropy for classification, squared error for regression). Gradient descent computes the slope of the loss with respect to every weight (via backpropagation in neural networks) and moves the weights against the slope by a step scaled by the learning rate. Epochs are full passes over the data; batches are the chunks processed per step. Regularisation (weight penalties, dropout, early stopping) fights overfitting; hyperparameters are the settings chosen by the engineer rather than learned (learning rate, layer sizes, tree depth).

Two practical variants from lecture 2: stochastic (mini-batch) gradient descent updates on a small batch of examples at a time, which is why training scales to millions of rows, and Newton's method takes far fewer but more expensive steps and suits small models such as logistic regression.
Basics refresher

A weight or parameter is one adjustable number inside the model. A gradient is a slope: which direction makes the error go down and how steeply. Underfitting is the opposite failure: the model is too simple to capture the pattern at all.

04

Learning curves tell you whether to buy more data or a better model

Plot error against the amount of training data. If training error is low and validation error is high, the model has memorised (high variance) and more data will close the gap. If both are high and close together, the model is too simple (high bias) and more data will not help. This one chart settles the most expensive argument in an ML roadmap.

High variance (overfitting) training examples β†’ error acceptable error validation error training error gap gap still closing β†’ more data will help High bias (underfitting) training examples β†’ error acceptable error validation error training error stuck above target both flat and high β†’ more data will not help
Andrew Ng's learning-curve diagnostic (CS229, lecture 8 and lecture 12). The x-axis is training set size, not training steps as in section 03. Solid blue is error on the training data, dashed orange is error on held-out data, dashed green is the error the product can live with.
What the curve showsDiagnosisWhat helpsWhat wastes money
Training error low, validation error high, gap still closingHigh varianceMore data, fewer or simpler features, regularisation, a smaller modelA bigger model
Both errors high, close together, flatHigh biasRicher features, a bigger or more flexible model, less regularisation, boostingMore data, more labelling
Both errors low but production results are poorDistribution mismatchData that looks like production; check drift and leakage (section 05)Any change to the model
Teaching notesAsk the learner to sketch the curve for a model their team is evaluating and say which regime it is in. Ng's rule for the class: if you cannot say whether the problem is bias or variance, you cannot say whether the next dollar goes to data or to modelling. The word 'bias' here is statistical, not the fairness sense.
Why a Head of Product cares
  • 'We need more data' is the default request and the most expensive one. A learning curve costs a day and tells you in advance whether the data will pay off.
  • High bias means the model cannot express the pattern, and no amount of labelling fixes that. Budget for features and modelling instead.
  • The diagnosis also sets the story for stakeholders. 'We are data-limited' and 'we are model-limited' lead to different quarters and different hires.
Go deeper

Bias is error from a model too simple to represent the truth; variance is error from sensitivity to the particular training sample. Expected error decomposes roughly into biasΒ², variance, and irreducible noise. Regularisation (CS229 lecture 8) trades variance for bias by penalising large weights, and cross-validation picks the strength. Ensembles map onto the same axes: bagging (random forests) averages many high-variance trees to cut variance, while boosting adds weak, high-bias learners in sequence to cut bias (lecture 9). Learning theory (the discussion section) gives the sample-complexity intuition: the examples needed to hold variance down grow roughly with the number of parameters, which is why a richer model asks for more data to reach the same validation error.

Basics refresher

A learning curve plots error against the number of training examples. Bias means the model is consistently wrong in the same way; variance means the model's answers would change a lot if you retrained it on a different sample of the same data. Regularisation is any penalty that keeps a model simpler than the data alone would make it.

05

The data set is split three ways, and both of the classic disasters are data problems, not model problems

Training data teaches, validation data tunes, and test data is opened once to report the honest score. Leakage (test information sneaking into training) inflates scores in the lab; drift (the world changing) deflates them in production.

One dataset, three jobs Train 70% the model learns from this Validate 15% tune settings Test 15% report once, at the end leakage: test information reaches training β†’ scores lie Drift: the world moves, the model does not training data today's inputs time β†’ retrain, or predictions decay
Top: the split and the one arrow that must never exist. Bottom: drift as two distributions; the model was fit to the solid one and is now scored on the dashed one.
Teaching notesGive a leakage example a PM will recognise: a churn model trained with a feature like 'account closed date', which is the answer in disguise. For drift, use a pricing model trained before a market shift.
Why a Head of Product cares
  • Ask who owns data quality, because nobody else in the loop can compensate for it.
  • Leakage is why lab results routinely beat production results. Insist on a time-based split for anything involving history.
  • Drift is the ongoing cost. A model with no monitoring is an unmonitored liability with a good launch deck.
Go deeper

Time-based splits (train on the past, test on the future) mimic production and expose leakage that random splits hide. Point-in-time correctness means every feature is computed using only data that existed at prediction time. Feature stores centralise those definitions so training and serving compute identical features (avoiding training-serving skew). Drift comes in kinds: input drift (feature distributions move), concept drift (the relationship between inputs and outcome moves), and label delay (you only learn the truth weeks later). Monitoring compares recent input distributions to the training distribution and alerts on divergence.

Split sizes follow data volume, not a fixed rule (lecture 8): 70/30 or 60/20/20 is sensible for thousands of examples, while with millions the validation and test sets can be 1% each, because they only need to be large enough to measure the difference you care about. With small data use k-fold cross-validation: split into k parts, train on k-1 and validate on the remaining one, rotate, and average, so every example does both jobs.
Basics refresher

A distribution describes how often each value appears; the bell curves above are distributions. A held-out set is data withheld from training. Ground truth is the real outcome, once it is known.

06

Six model families cover almost everything, and the boring ones win on tabular data

Gradient-boosted trees still beat deep learning on rows-and-columns problems, and a logistic regression is often good enough. Neural networks and transformers earn their cost on images, audio, and language.

harder to explain β†’ capacity ↑ Linear / logistic regression baseline for tabular data Decision tree rules you can read Random forest many trees, averaged Gradient boosting wins most tabular contests Neural network images, audio, sequences Transformer / LLM language, multimodal For rows-and-columns problems, start at the left; move right only when the data is unstructured
Position is approximate and the axes trade off: every step right buys capacity and spends explainability and cost. The table below is the same information as a decision aid.
FamilyBest forWatch out forExplainable?
Linear / logistic regressionBaselines, pricing, risk scores with few featuresMisses interactions and non-linear patternsYes, coefficient by coefficient
Decision treeRules people must read and approveUnstable; overfits aloneYes, as a flowchart
Random forestRobust default for tabular dataLarge, slower to servePartly (feature importance)
Gradient boosting (XGBoost, LightGBM)Most tabular prediction: fraud, churn, demandNeeds tuning; can overfitPartly (SHAP values)
Neural networkImages, audio, sequences, very large dataData hungry, GPU cost, opaqueWeakly
Transformer / LLMLanguage, code, multimodal, generationCost, latency, hallucinationWeakly, but can explain in words
Teaching notesThe message for leaders is that 'use AI' does not mean 'use an LLM'. Have the learner pick a family for churn prediction (boosting), for a lending decision that regulators review (logistic or tree), and for support-ticket summarisation (LLM).
Why a Head of Product cares
  • Choosing a simpler family is a legitimate product decision when explainability, latency, or cost matter more than the last point of accuracy.
  • Regulated decisions (credit, hiring, health) often require explainable models; this constrains you to the left half.
  • Ask for the baseline. A team that cannot beat logistic regression by a meaningful margin should ship logistic regression.
Go deeper

Ensembles combine many weak models: forests average trees trained on random subsets; boosting trains trees sequentially on the previous errors. SHAP values attribute a prediction to features, giving partial explainability to ensembles. Neural networks generalise the linear model by stacking layers with non-linear activations; convolutional layers suit images, recurrent layers suited sequences until transformers replaced them. Choosing among families is done empirically with cross-validation on the validation set; the winner is rarely predictable in advance, which is why teams try several.

The two ensemble styles fix different failures (lecture 9): bagging cuts variance, boosting cuts bias; section 04 explains which one your model has. Support vector machines with kernels (lectures 6 and 7) were the state of the art before deep learning and remain strong on small, high-dimensional problems; a kernel lets the model work from similarity between examples instead of raw features, the same idea embeddings use in section 12. Non-parametric methods such as nearest neighbours and locally weighted regression (lecture 3) keep the training data around at prediction time, so serving cost grows with the data.
Basics refresher

Tabular data is a spreadsheet: one row per example, one column per feature. Interpretability is whether a person can see why the model produced a given output. A baseline is the simplest reasonable model, used to judge whether a complex one is worth it.

07

A neural network is layers of weighted sums, and depth is what makes it powerful

Each node multiplies its inputs by learned weights, adds them up, and passes the result through a simple non-linear function. Stack enough layers and the network can represent almost any pattern, which is why the same idea recognises faces, transcribes speech, and writes text.

Inputs Hidden layer Hidden layer Outputs amount hour country fraud legit each line is a weight, a learned number; each node adds up its inputs and applies a simple non-linear squash this toy has 46 weights; a frontier LLM has hundreds of billions, which is why they run on GPUs
Data flows left to right; training flows right to left (errors propagate back to adjust the weights). The number of weights is the 'parameters' figure vendors quote.
Teaching notesCount the lines with the learner to make 'parameters' concrete, then scale up: a 70-billion-parameter model is this picture with a billion and a half times more lines. That is the reason for GPUs and for inference cost.
Why a Head of Product cares
  • Parameter count is a rough proxy for capability and a direct driver of serving cost. More is not automatically better for your task.
  • Deep learning needs a lot of data and compute to beat simpler models; for small tabular problems it usually does not.
  • GPU capacity is a supply-chain and budget item for any team hosting its own models.
Go deeper

Each layer computes output = activation(W Γ— input + b). Activations (ReLU, GELU) introduce non-linearity; without them, stacked layers collapse to one linear model. Backpropagation applies the chain rule to compute gradients for every weight efficiently. Specialised layers encode structure: convolutions share weights across positions in an image; attention (Module 1, section 3) lets each position in a sequence weigh all others. Mixture-of-experts models route each token through a subset of the network, so a very large model can be cheap per token.

Lecture 11 covers the fixes that make deep networks trainable in practice: careful weight initialisation and normalisation so signals neither vanish nor explode across many layers, mini-batch training, and adaptive optimisers such as Adam. These are what a data scientist means by 'the training was unstable'.
Basics refresher

A GPU is a chip built to run thousands of small multiplications in parallel, originally for graphics, now for the matrix arithmetic in this picture. Parameters and weights mean the same thing here. A layer is one column of nodes.

08

Accuracy hides the two mistakes that matter; precision and recall name them

A model that never flags fraud is 99% accurate if 1% of transactions are fraud. Precision and recall separate 'when we act, are we right' from 'do we catch what matters', and the threshold between them is a business decision, not a technical one.

Confusion matrix (fraud example) predicted: fraud predicted: legit actual fraud actual legit True positive fraud caught False negative fraud missed False positive good customer blocked True negative correctly ignored Two questions, two numbers Precision = TP Γ· (TP + FP) when we flag something, how often are we right? Recall = TP Γ· (TP + FN) of all the real fraud, how much did we catch? strict threshold loose threshold ↑ precision, ↓ recall ↓ precision, ↑ recall the threshold is a business decision: which mistake costs more?
Left: every prediction lands in one of four cells; the two red cells are the two kinds of mistake. Right: sliding the decision threshold trades one mistake for the other.
Teaching notesHave the learner assign a dollar cost to each red cell for their own use case. The threshold falls out of that arithmetic. Then point out that the same model can be 'precise' or 'thorough' depending on where the product team sets it.
Why a Head of Product cares
  • Own the threshold. It encodes your tolerance for each mistake and should be set by the people who understand the cost of each.
  • Ask for precision and recall (or the full matrix) rather than accuracy, especially when the positive class is rare.
  • The two mistakes usually have different owners: false positives hit customer experience, false negatives hit risk or revenue. Make both visible in the dashboard.
Go deeper

F1 is the harmonic mean of precision and recall, a single number when you need one. The ROC curve plots true-positive rate against false-positive rate across thresholds; AUC summarises it (0.5 is coin-flip, 1.0 is perfect). For imbalanced problems the precision-recall curve is more informative than ROC. Regression models use error metrics (MAE, RMSE) and should be compared to a naive baseline (predict the average, or yesterday's value). Calibration asks whether a predicted 80% actually happens 80% of the time; it matters when the score itself is shown to users or used in expected-value calculations.

Basics refresher

A threshold is the score above which the model says 'yes'. A positive is the thing you are looking for (fraud, churn, a match). A rare class is a positive that appears in a small fraction of examples, which is where accuracy misleads.

09

Before you improve the model, find out which part of it is failing

Build the quick-and-dirty version first, then diagnose. Error analysis swaps each pipeline stage for a perfect version to see where accuracy jumps; ablative analysis removes each improvement to see which ones earned their keep. Both stop a team spending a quarter on the wrong component.

Clean text strip signatures, quotes Detect language en, de, fr… Extract entities product, order, plan Classify intent refund, bug, how-to Route to queue rules on the above 80% 90% 100% today: 85% +0.5 85.5% +1 86.5% +4.5 91% +7 ← fix this first 98% +2 100% whole-system accuracy after swapping in a perfect, hand-labelled output for each stage, cumulatively from the left
Error analysis on a pipeline, Ng's method from CS229 lecture 12 applied to ticket routing. Each bar replaces one more stage with ground truth. The biggest jump names the bottleneck; a perfect language detector is worth one point, so a language-detection project is not worth funding.
Ablative analysis: remove one component at a timeAccuracyWhat it earned
Full anti-spam system (Ng's example from the CS229 advice slides)99.9%
Without spelling correction99.0%0.9 points
Without sender-host features98.9%0.1
Without email-header features98.9%0.0
Without text-parser features98.0%0.9
Without the JavaScript parser94.5%3.5
Without image features (back to the plain baseline)94.0%0.5
Teaching notesGive the learner a pipeline from their own product and ask which stage they would bet is the bottleneck. Then ask how they would test the bet in a week without building anything: label 100 to 200 examples per stage and swap them in. Most people bet on the wrong stage, which is the point.
Why a Head of Product cares
  • Diagnosis is cheap and modelling is expensive. Ng's order of work: quick-and-dirty system first, error analysis second, investment third.
  • Error analysis gives a ceiling for each component. If a perfect stage moves the system by half a point, no project on that stage is worth funding.
  • Ablative analysis shows what a team's last six months actually contributed, and what can be deleted. Header features above earned nothing and still cost maintenance.
  • The most common finding is that the metric was wrong: the system optimised the number it was given and the number was not the product. See the helicopter test below.
Go deeper

Ng's helicopter diagnostic separates two failures that look identical from outside: the optimiser did not find a good solution, or it found a solution that is good by your objective and bad in reality. The test is to score a human expert's behaviour with your own cost function. If the human scores better than the model, the optimiser is failing (train longer, change the algorithm). If the model scores better than the human but flies worse, the objective is wrong and the fix is the metric, not the model. The same test applies to LLM evals and reward models: check whether the answer a human prefers scores higher under your rubric. Error analysis needs ground truth for intermediate stages, which means labelling a small sample per stage; 100 to 200 examples usually suffice. Read the misclassified examples by hand and tally them by cause; the histogram of causes is the roadmap. Lecture 20 applies the same three-way split to reinforcement learning: is it the simulator, the algorithm, or the reward function?

Basics refresher

A pipeline stage is one step that consumes the previous step's output. Ablation means removing one part and measuring the difference. Ground truth for a stage is the output a careful human would have produced at that step. An objective or cost function is the number the training process is told to optimise.

10

Shipping a model is shipping software plus a second pipeline for data and weights

The model pipeline mirrors the code pipeline: data instead of commits, training instead of building, evaluation instead of tests, a registry instead of a release, and a retrain loop that code does not need. Teams that already run CI/CD well are most of the way there.

Software delivery (CI/CD) Commit Build Test Deploy Monitor Model delivery (MLOps) Data Train Evaluate Register Serve monitor β†’ drift detected β†’ retrain with fresh data same tools, same gates
The top lane is ByteByteGo's CI/CD crash course in one line. The bottom lane adds what ML needs: versioned data, a model registry, and monitoring that can trigger retraining.
Teaching notesAsk what 'test' means for a model (an eval on held-out data plus checks on latency and fairness) and what 'rollback' means (the registry keeps the previous version). Then ask who is on call for a model that drifts at 2 a.m.
Why a Head of Product cares
  • The infrastructure investment is mostly reusable: pipelines, registries, and monitoring serve every model after the first.
  • Gates matter as much as speed. An eval that must pass before a model is registered is the ML equivalent of a failing test blocking a merge.
  • Plan the on-call. A model in production has failure modes (drift, latency, upstream data breaks) that need an owner.
Go deeper

Core components: a feature store (consistent feature definitions for training and serving), an experiment tracker (which data, code, and hyperparameters produced which metrics), a model registry (versioned artefacts with stage: staging, production, archived), a serving layer (batch scoring or a real-time endpoint), and monitoring (input drift, prediction drift, latency, and outcome metrics when labels arrive). Rollouts use shadow mode (score but do not act), canaries, and A/B tests. From the CI/CD source: continuous integration automates build and test on every commit; continuous delivery keeps code deployable at all times; continuous deployment pushes every passing change to production automatically.

Basics refresher

CI/CD is continuous integration and continuous delivery or deployment: automation that builds, tests, and ships code on every change. A registry is a versioned store of built artefacts. On-call is the rota of people paged when production breaks.

11

Three questions decide whether you need an LLM at all

If the input is tabular, use classical ML. If it is language but the task is narrow and high-volume, a small tuned model is cheaper and faster. Reserve LLMs for open-ended work that needs judgement, and prototype with an LLM before deciding to specialise.

Is the input mostly numbers in rows and columns? yes Classical ML boosting, regression no text, images, audio Is the task narrow and repeated millions of times? yes Small tuned model cheap, fast, specific no open-ended, needs judgement LLM prompt, RAG, or agent Fraud scoring, demand forecasts, churn, pricing Ticket routing, sentiment, entity extraction at scale Drafting, summarising, answering, multi-step tasks
The flow is a default, not a law. A common pattern is to prototype with an LLM to learn the task, then distil the narrow, high-volume part into a small model.
Teaching notesRun the learner's backlog through the three questions. Most items resolve in under a minute, and the exercise usually reveals one 'AI feature' that is really a rules feature.
Why a Head of Product cares
  • LLM cost at scale is the number that surprises finance. A narrow task at a million calls a day almost always wants a small model.
  • Prototype with the LLM anyway: it is the fastest way to learn what the task really needs before investing in a specialised model.
  • Mixed architectures are normal: an LLM for the conversation, classical models for the scores it consults.
Go deeper

LLMs can act as zero-shot classifiers, which makes them ideal for prototyping and for labelling the data a small model later trains on (LLM-assisted labelling). Small language models (under ~10B parameters) fine-tuned on that data often match the LLM on the narrow task at a fraction of latency and cost (ByteByteGo's LLM vs SLM comparison). Embedding models plus a classifier are another cheap middle path for text. The decision should be revisited as prices fall.

Basics refresher

A small language model is an LLM-style model with far fewer parameters, often run on modest hardware. Zero-shot means the model performs a task from an instruction alone with no examples. Entity extraction pulls names, dates, and amounts out of text.

12

Embeddings turn meaning into coordinates, which turns search into geometry

An embedding model maps text (or images, or products) to a point in a space with thousands of dimensions, arranged so that similar meanings are close together. Finding relevant documents becomes finding nearest neighbours, and that is the engine under RAG, semantic search, and recommendations.

dimension 1 (of thousands) dim 2 king queen prince throne coffee tea latte espresso invoice receipt payment 'hot drink' nearest neighbours = the retrieval step in RAG similar meaning β†’ nearby points
A two-dimensional cartoon of a space that really has a thousand or more dimensions. The dashed circle is a nearest-neighbour query: the retrieval step from Module 1's RAG diagram.
Teaching notesConnect this back to RAG explicitly: the vector store holds these points for every passage, and a question becomes a point whose neighbours are the passages to show the model. Ask what happens when two documents say the same thing in different words (they land close together, which keyword search misses).
Why a Head of Product cares
  • Embeddings are cheap and reusable: one embedding pipeline powers search, deduplication, clustering, and recommendations.
  • Retrieval quality is mostly embedding quality plus chunking. Evaluate retrieval separately from generation.
  • Embedding models are swappable, but switching one means re-embedding everything; plan the migration.
Go deeper

Embedding models are trained so that texts with similar meaning have high cosine similarity. Approximate nearest-neighbour indexes (HNSW and similar) make search fast at millions of vectors. Embeddings of different types can share a space (CLIP puts images and captions together), enabling image search by text. Hybrid search combines vectors with keyword scoring so exact terms (product codes, names) still match. Evaluate with recall@k: for known question-passage pairs, how often is the right passage in the top k results?

Basics refresher

A dimension is one axis of a space; the picture has two, real embeddings have hundreds or thousands. Nearest neighbour means the stored point with the smallest distance to the query point. Semantic search finds by meaning rather than by matching words.

13

Inference cost is a set of knobs, and half of them cost you no quality at all

Smaller, quantised, or distilled models trade a little quality for a lot of speed and cost. Caching, batching, and speculative decoding keep outputs identical and still cut the bill. Know which knob your team turned before you accept a quality regression.

Quality Latency Cost Use a smaller model ↓ slightly ↓↓ ↓↓ Quantise weights (8- or 4-bit) ↓ tiny ↓ ↓ Distil into a student model ↓ some ↓↓ ↓↓ Cache repeated prompt prefixes = ↓ ↓ Batch many requests = ↑ per request ↓↓ Speculative decoding = ↓ = green = improves for you, red = worsens, grey = unchanged; the first three change outputs, the last three do not
From the Google LLM whitepaper's inference chapter: output-approximating techniques (top three) versus output-preserving ones (bottom three). Batching lowers cost per token by raising per-request latency, which suits offline jobs and not chat.
Teaching notesAsk which knobs are free (the bottom three) and why teams do not always use them (engineering effort, and batching hurts interactive latency). Then ask for the task-level eval before and after any of the top three.
Why a Head of Product cares
  • Ask for output-preserving optimisations first; they are pure savings.
  • A quality drop from a smaller or quantised model is only real if your eval set says so. Benchmarks are not your task.
  • Interactive and batch workloads want different settings; separate them so batch jobs do not slow the chat product.
Go deeper

Prefix caching stores the computed attention state for a shared prompt prefix (your system prompt and few-shot examples) so only the new part is processed. Continuous batching fills the GPU with many sequences at different stages. Speculative decoding lets a small draft model propose several tokens that the large model verifies in one pass; outputs are identical to the large model alone. Flash attention reorganises the attention computation to use GPU memory more efficiently. Quantisation-aware training recovers most of the quality lost by low precision. Whitepaper example: a 2Γ— speed-up for a 2% accuracy drop on a detection task, an acceptable trade in most products.

Basics refresher

Precision here means how many bits store each number (32-bit float vs 8-bit integer), not the classification metric. Cache means keeping a computed result to reuse instead of recomputing. Batch means processing many items together.

Glossary

FeatureOne input variable to a model.
LabelThe known correct answer for a training example.
LossA number measuring how wrong the model is; training minimises it.
Gradient descentThe step-downhill procedure that adjusts weights during training.
OverfittingMemorising the training data instead of learning the pattern.
Validation setData used to tune settings; test set is opened once at the end.
LeakageInformation from the answer or the future sneaking into training.
DriftProduction inputs or relationships moving away from the training data.
PrecisionOf the things we flagged, the share that were right.
RecallOf the real positives, the share we caught.
ThresholdThe score above which the model says yes; a business decision.
Gradient boostingSequential trees that correct each other; the tabular workhorse.
Neural networkLayers of weighted sums with non-linear activations.
EmbeddingA vector that encodes meaning so similar things are close.
Model registryVersioned store of trained models with their stage and metrics.
QuantisationStoring weights in fewer bits to cut memory and cost.
Bias (statistical)Error from a model too simple to capture the pattern; more data does not fix it.
VarianceError from sensitivity to the training sample; more data or regularisation fixes it.
Learning curveError plotted against training set size; the bias-vs-variance diagnostic.
Cross-validationRotating which slice of the data is held out, so small data can both train and validate.
Error analysisSwapping each pipeline stage for ground truth to find the stage that caps accuracy.
Ablative analysisRemoving components one at a time to measure what each contributed.
Generative modelLearns what each class looks like (Naive Bayes); not the same sense as generative AI.
Reward functionThe score a reinforcement-learning agent optimises; in product terms, the spec.

Questions to ask your team

  1. What exactly is the target variable, and what decision does the prediction feed?
  2. Where do the labels come from, how many do we have, and what does each new one cost?
  3. Was the test set separated in time from training? Has anyone tuned against it?
  4. What is the baseline, and by how much does the proposed model beat it?
  5. Are we bias-limited or variance-limited? Show me the learning curve before we buy more data.
  6. Which pipeline stage caps our accuracy? Have we swapped in ground truth for each stage and measured?
  7. If a human expert's output scores worse on our metric than the model's does, is the metric wrong?
  8. Show me the confusion matrix. What does each red cell cost us in dollars?
  9. Who sets the threshold, and when was it last reviewed?
  10. What is monitored in production, and what triggers a retrain?
  11. Can we roll back to the previous model version in under an hour?
  12. For the LLM parts: which knobs have we turned, and did evals change?

Reading list

  • Stanford CS229: Machine Learning, Andrew Ng (Autumn 2018) β€” YouTube playlist, 21 lectures of about 80 minutes each. The canonical university course under this module. For a product leader the three to watch in full are Lecture 1 (framing), Lecture 8 (bias, variance, data splits) and Lecture 12 (debugging and error analysis); the map below points the rest at the section they support. Lectures 13 to 19 (EM, factor analysis, ICA, control theory) are beyond what this module needs.
  • ByteByteGo (Alex Xu) β€” A Crash Course in CI/CD (the top lane of section 10); Good Code vs. Bad Code (why the code around a model matters as much as the model); Understanding Database Types (where vector stores fit); Large vs Small Language Models (section 11's middle branch); How to Shrink a Language Model Without Making it Too Dumb (quantisation and distillation, section 13); How Large Language Models Learn (training, visually).
  • Foundational Large Language Models & Text Generation (Google, 2025) β€” kaggle.com. Sections on training, fine-tuning, and the inference trade-offs behind section 13.
Module sectionCS229 lectureTake from it
01 Three kinds of learningLecture 1, Lecture 5, Lecture 16Mitchell's T-E-P definition; generative vs discriminative; what an MDP is
03 What training isLecture 2, Lecture 3Gradient descent on linear and logistic regression, batch vs stochastic
04 Bias vs varianceLecture 8, Learning theory sectionLearning curves, regularisation, why parameters demand data
05 Data is the productLecture 8Split sizes by data volume, k-fold cross-validation, feature selection
06 The model zooLecture 9, Lecture 6, Lecture 7Trees, bagging vs boosting; SVMs and kernels
07 Neural networksLecture 10, Lecture 11Layers and backprop; initialisation, normalisation, Adam
09 Diagnosing failuresLecture 12, Lecture 20Error analysis, ablation, optimiser-vs-objective; RL diagnostics
Unsupervised toolkit (in 01)Lecture 13, Lecture 15k-means and mixtures; PCA