Machine learning comes in three flavours, defined by what feedback the model gets
Supervised learning has the right answers; unsupervised learning has only the data; reinforcement learning has a reward. Almost every business use case is supervised, which means someone has to produce the labels.
Why a Head of Product cares
- The labelling budget is the real cost of supervised ML. Historical outcomes (did the customer churn?) are free labels; human judgements (is this review toxic?) are not.
- Unsupervised results need a person to interpret them; plan for analyst time, not just model time.
- Reinforcement learning is powerful but rare in products because it needs a safe place to try things. Its main mainstream use today is shaping LLMs.
Go deeper
Supervised tasks split into classification (discrete label) and regression (a number). Unsupervised methods include clustering (k-means), dimensionality reduction (PCA, UMAP), and anomaly detection. Self-supervised learning, where the label is derived from the data itself (predict the next word, fill the blank), is how LLMs are pretrained and blurs the first two categories. Semi-supervised and active learning reduce labelling cost by choosing which examples to label.
Generative vs discriminative is a second split inside supervised learning (CS229 lecture 5): discriminative models (logistic regression, trees, most neural networks) learn the boundary between classes directly, while generative models (Naive Bayes, Gaussian discriminant analysis) learn what each class looks like and ask which one a new input resembles. Generative models need less data and were the workhorse of spam filtering; discriminative models win once data is plentiful. This is a different sense of 'generative' from 'generative AI', which just means a model that produces content. Reinforcement learning is formalised as a Markov decision process: states, actions, a reward, and the rules for how actions change the world (lectures 16 to 20). The product decision is the reward function: it is the spec, and an agent optimises exactly what it says rather than what you meant. The unsupervised toolkit in the middle panel is k-means and Gaussian mixtures for grouping (lecture 13) and PCA for compressing many features into a few (lecture 15).Basics refresher
A label is the known outcome for a training example. A feature is one input variable (age, amount, number of logins). A prediction is the model's output for a new example.
Tom Mitchell's definition, used in CS229 lecture 1: a program learns from experience E at task T if its performance P on T improves with E. Every ML project can be written as that triple, and if one of the three is missing the project is not ready.An ML system is a loop, not a launch
Framing, data, training, evaluation, deployment, and monitoring feed each other. The model you ship is a snapshot; the loop is the asset, and the team that owns the loop owns the outcome.
Why a Head of Product cares
- Budget for the loop, not the model: data pipelines, labelling, evaluation, and monitoring are recurring costs.
- Framing decides everything downstream. 'Predict churn' is not a frame; 'flag accounts to call this week, at most 200, ranked by expected saved revenue' is.
- Retraining cadence is a product decision with a cost. Daily for fraud, quarterly for demand forecasting, when-evidence-says for most things.
Go deeper
Framing includes choosing the target variable (what exactly is being predicted), the unit of prediction, the decision the prediction feeds, and the metric that reflects business cost. Data work includes joining sources, handling missing values, deduplication, and defining the point in time at which features are known (to avoid leakage). Training compares several model families with cross-validation. Evaluation is done on a held-out test set and then in a shadow or A/B deployment. Monitoring tracks input drift, prediction drift, and, where available, realised outcomes.
Basics refresher
A pipeline is a chain of automated steps that turn raw data into a trained model or a prediction. Deployment means the model runs on a server that other software can call. An API endpoint is the address that software calls.
Training is rolling downhill on an error landscape, one small step at a time
A model starts with random weights. Training measures how wrong it is on the examples (the loss), then nudges every weight a little in the direction that reduces the error, and repeats millions of times. Stop too late and it memorises the examples instead of learning the pattern.
Why a Head of Product cares
- 'It scored 99% in testing' means nothing until you know which data it was tested on. Ask for held-out results.
- Training cost is compute time multiplied by iterations; for deep learning that is GPU hours, and for LLMs it is the reason only labs pretrain.
- Overfitting is the ML version of a demo that works only on the demo data. The defence is a validation set the team never tunes against.
Go deeper
The loss function turns predictions and labels into one number (cross-entropy for classification, squared error for regression). Gradient descent computes the slope of the loss with respect to every weight (via backpropagation in neural networks) and moves the weights against the slope by a step scaled by the learning rate. Epochs are full passes over the data; batches are the chunks processed per step. Regularisation (weight penalties, dropout, early stopping) fights overfitting; hyperparameters are the settings chosen by the engineer rather than learned (learning rate, layer sizes, tree depth).
Two practical variants from lecture 2: stochastic (mini-batch) gradient descent updates on a small batch of examples at a time, which is why training scales to millions of rows, and Newton's method takes far fewer but more expensive steps and suits small models such as logistic regression.Basics refresher
A weight or parameter is one adjustable number inside the model. A gradient is a slope: which direction makes the error go down and how steeply. Underfitting is the opposite failure: the model is too simple to capture the pattern at all.
Learning curves tell you whether to buy more data or a better model
Plot error against the amount of training data. If training error is low and validation error is high, the model has memorised (high variance) and more data will close the gap. If both are high and close together, the model is too simple (high bias) and more data will not help. This one chart settles the most expensive argument in an ML roadmap.
| What the curve shows | Diagnosis | What helps | What wastes money |
|---|---|---|---|
| Training error low, validation error high, gap still closing | High variance | More data, fewer or simpler features, regularisation, a smaller model | A bigger model |
| Both errors high, close together, flat | High bias | Richer features, a bigger or more flexible model, less regularisation, boosting | More data, more labelling |
| Both errors low but production results are poor | Distribution mismatch | Data that looks like production; check drift and leakage (section 05) | Any change to the model |
Why a Head of Product cares
- 'We need more data' is the default request and the most expensive one. A learning curve costs a day and tells you in advance whether the data will pay off.
- High bias means the model cannot express the pattern, and no amount of labelling fixes that. Budget for features and modelling instead.
- The diagnosis also sets the story for stakeholders. 'We are data-limited' and 'we are model-limited' lead to different quarters and different hires.
Go deeper
Bias is error from a model too simple to represent the truth; variance is error from sensitivity to the particular training sample. Expected error decomposes roughly into biasΒ², variance, and irreducible noise. Regularisation (CS229 lecture 8) trades variance for bias by penalising large weights, and cross-validation picks the strength. Ensembles map onto the same axes: bagging (random forests) averages many high-variance trees to cut variance, while boosting adds weak, high-bias learners in sequence to cut bias (lecture 9). Learning theory (the discussion section) gives the sample-complexity intuition: the examples needed to hold variance down grow roughly with the number of parameters, which is why a richer model asks for more data to reach the same validation error.
Basics refresher
A learning curve plots error against the number of training examples. Bias means the model is consistently wrong in the same way; variance means the model's answers would change a lot if you retrained it on a different sample of the same data. Regularisation is any penalty that keeps a model simpler than the data alone would make it.
The data set is split three ways, and both of the classic disasters are data problems, not model problems
Training data teaches, validation data tunes, and test data is opened once to report the honest score. Leakage (test information sneaking into training) inflates scores in the lab; drift (the world changing) deflates them in production.
Why a Head of Product cares
- Ask who owns data quality, because nobody else in the loop can compensate for it.
- Leakage is why lab results routinely beat production results. Insist on a time-based split for anything involving history.
- Drift is the ongoing cost. A model with no monitoring is an unmonitored liability with a good launch deck.
Go deeper
Time-based splits (train on the past, test on the future) mimic production and expose leakage that random splits hide. Point-in-time correctness means every feature is computed using only data that existed at prediction time. Feature stores centralise those definitions so training and serving compute identical features (avoiding training-serving skew). Drift comes in kinds: input drift (feature distributions move), concept drift (the relationship between inputs and outcome moves), and label delay (you only learn the truth weeks later). Monitoring compares recent input distributions to the training distribution and alerts on divergence.
Split sizes follow data volume, not a fixed rule (lecture 8): 70/30 or 60/20/20 is sensible for thousands of examples, while with millions the validation and test sets can be 1% each, because they only need to be large enough to measure the difference you care about. With small data use k-fold cross-validation: split into k parts, train on k-1 and validate on the remaining one, rotate, and average, so every example does both jobs.Basics refresher
A distribution describes how often each value appears; the bell curves above are distributions. A held-out set is data withheld from training. Ground truth is the real outcome, once it is known.
Six model families cover almost everything, and the boring ones win on tabular data
Gradient-boosted trees still beat deep learning on rows-and-columns problems, and a logistic regression is often good enough. Neural networks and transformers earn their cost on images, audio, and language.
| Family | Best for | Watch out for | Explainable? |
|---|---|---|---|
| Linear / logistic regression | Baselines, pricing, risk scores with few features | Misses interactions and non-linear patterns | Yes, coefficient by coefficient |
| Decision tree | Rules people must read and approve | Unstable; overfits alone | Yes, as a flowchart |
| Random forest | Robust default for tabular data | Large, slower to serve | Partly (feature importance) |
| Gradient boosting (XGBoost, LightGBM) | Most tabular prediction: fraud, churn, demand | Needs tuning; can overfit | Partly (SHAP values) |
| Neural network | Images, audio, sequences, very large data | Data hungry, GPU cost, opaque | Weakly |
| Transformer / LLM | Language, code, multimodal, generation | Cost, latency, hallucination | Weakly, but can explain in words |
Why a Head of Product cares
- Choosing a simpler family is a legitimate product decision when explainability, latency, or cost matter more than the last point of accuracy.
- Regulated decisions (credit, hiring, health) often require explainable models; this constrains you to the left half.
- Ask for the baseline. A team that cannot beat logistic regression by a meaningful margin should ship logistic regression.
Go deeper
Ensembles combine many weak models: forests average trees trained on random subsets; boosting trains trees sequentially on the previous errors. SHAP values attribute a prediction to features, giving partial explainability to ensembles. Neural networks generalise the linear model by stacking layers with non-linear activations; convolutional layers suit images, recurrent layers suited sequences until transformers replaced them. Choosing among families is done empirically with cross-validation on the validation set; the winner is rarely predictable in advance, which is why teams try several.
The two ensemble styles fix different failures (lecture 9): bagging cuts variance, boosting cuts bias; section 04 explains which one your model has. Support vector machines with kernels (lectures 6 and 7) were the state of the art before deep learning and remain strong on small, high-dimensional problems; a kernel lets the model work from similarity between examples instead of raw features, the same idea embeddings use in section 12. Non-parametric methods such as nearest neighbours and locally weighted regression (lecture 3) keep the training data around at prediction time, so serving cost grows with the data.Basics refresher
Tabular data is a spreadsheet: one row per example, one column per feature. Interpretability is whether a person can see why the model produced a given output. A baseline is the simplest reasonable model, used to judge whether a complex one is worth it.
A neural network is layers of weighted sums, and depth is what makes it powerful
Each node multiplies its inputs by learned weights, adds them up, and passes the result through a simple non-linear function. Stack enough layers and the network can represent almost any pattern, which is why the same idea recognises faces, transcribes speech, and writes text.
Why a Head of Product cares
- Parameter count is a rough proxy for capability and a direct driver of serving cost. More is not automatically better for your task.
- Deep learning needs a lot of data and compute to beat simpler models; for small tabular problems it usually does not.
- GPU capacity is a supply-chain and budget item for any team hosting its own models.
Go deeper
Each layer computes output = activation(W Γ input + b). Activations (ReLU, GELU) introduce non-linearity; without them, stacked layers collapse to one linear model. Backpropagation applies the chain rule to compute gradients for every weight efficiently. Specialised layers encode structure: convolutions share weights across positions in an image; attention (Module 1, section 3) lets each position in a sequence weigh all others. Mixture-of-experts models route each token through a subset of the network, so a very large model can be cheap per token.
Lecture 11 covers the fixes that make deep networks trainable in practice: careful weight initialisation and normalisation so signals neither vanish nor explode across many layers, mini-batch training, and adaptive optimisers such as Adam. These are what a data scientist means by 'the training was unstable'.Basics refresher
A GPU is a chip built to run thousands of small multiplications in parallel, originally for graphics, now for the matrix arithmetic in this picture. Parameters and weights mean the same thing here. A layer is one column of nodes.
Accuracy hides the two mistakes that matter; precision and recall name them
A model that never flags fraud is 99% accurate if 1% of transactions are fraud. Precision and recall separate 'when we act, are we right' from 'do we catch what matters', and the threshold between them is a business decision, not a technical one.
Why a Head of Product cares
- Own the threshold. It encodes your tolerance for each mistake and should be set by the people who understand the cost of each.
- Ask for precision and recall (or the full matrix) rather than accuracy, especially when the positive class is rare.
- The two mistakes usually have different owners: false positives hit customer experience, false negatives hit risk or revenue. Make both visible in the dashboard.
Go deeper
F1 is the harmonic mean of precision and recall, a single number when you need one. The ROC curve plots true-positive rate against false-positive rate across thresholds; AUC summarises it (0.5 is coin-flip, 1.0 is perfect). For imbalanced problems the precision-recall curve is more informative than ROC. Regression models use error metrics (MAE, RMSE) and should be compared to a naive baseline (predict the average, or yesterday's value). Calibration asks whether a predicted 80% actually happens 80% of the time; it matters when the score itself is shown to users or used in expected-value calculations.
Basics refresher
A threshold is the score above which the model says 'yes'. A positive is the thing you are looking for (fraud, churn, a match). A rare class is a positive that appears in a small fraction of examples, which is where accuracy misleads.
Before you improve the model, find out which part of it is failing
Build the quick-and-dirty version first, then diagnose. Error analysis swaps each pipeline stage for a perfect version to see where accuracy jumps; ablative analysis removes each improvement to see which ones earned their keep. Both stop a team spending a quarter on the wrong component.
| Ablative analysis: remove one component at a time | Accuracy | What it earned |
|---|---|---|
| Full anti-spam system (Ng's example from the CS229 advice slides) | 99.9% | |
| Without spelling correction | 99.0% | 0.9 points |
| Without sender-host features | 98.9% | 0.1 |
| Without email-header features | 98.9% | 0.0 |
| Without text-parser features | 98.0% | 0.9 |
| Without the JavaScript parser | 94.5% | 3.5 |
| Without image features (back to the plain baseline) | 94.0% | 0.5 |
Why a Head of Product cares
- Diagnosis is cheap and modelling is expensive. Ng's order of work: quick-and-dirty system first, error analysis second, investment third.
- Error analysis gives a ceiling for each component. If a perfect stage moves the system by half a point, no project on that stage is worth funding.
- Ablative analysis shows what a team's last six months actually contributed, and what can be deleted. Header features above earned nothing and still cost maintenance.
- The most common finding is that the metric was wrong: the system optimised the number it was given and the number was not the product. See the helicopter test below.
Go deeper
Ng's helicopter diagnostic separates two failures that look identical from outside: the optimiser did not find a good solution, or it found a solution that is good by your objective and bad in reality. The test is to score a human expert's behaviour with your own cost function. If the human scores better than the model, the optimiser is failing (train longer, change the algorithm). If the model scores better than the human but flies worse, the objective is wrong and the fix is the metric, not the model. The same test applies to LLM evals and reward models: check whether the answer a human prefers scores higher under your rubric. Error analysis needs ground truth for intermediate stages, which means labelling a small sample per stage; 100 to 200 examples usually suffice. Read the misclassified examples by hand and tally them by cause; the histogram of causes is the roadmap. Lecture 20 applies the same three-way split to reinforcement learning: is it the simulator, the algorithm, or the reward function?
Basics refresher
A pipeline stage is one step that consumes the previous step's output. Ablation means removing one part and measuring the difference. Ground truth for a stage is the output a careful human would have produced at that step. An objective or cost function is the number the training process is told to optimise.
Shipping a model is shipping software plus a second pipeline for data and weights
The model pipeline mirrors the code pipeline: data instead of commits, training instead of building, evaluation instead of tests, a registry instead of a release, and a retrain loop that code does not need. Teams that already run CI/CD well are most of the way there.
Why a Head of Product cares
- The infrastructure investment is mostly reusable: pipelines, registries, and monitoring serve every model after the first.
- Gates matter as much as speed. An eval that must pass before a model is registered is the ML equivalent of a failing test blocking a merge.
- Plan the on-call. A model in production has failure modes (drift, latency, upstream data breaks) that need an owner.
Go deeper
Core components: a feature store (consistent feature definitions for training and serving), an experiment tracker (which data, code, and hyperparameters produced which metrics), a model registry (versioned artefacts with stage: staging, production, archived), a serving layer (batch scoring or a real-time endpoint), and monitoring (input drift, prediction drift, latency, and outcome metrics when labels arrive). Rollouts use shadow mode (score but do not act), canaries, and A/B tests. From the CI/CD source: continuous integration automates build and test on every commit; continuous delivery keeps code deployable at all times; continuous deployment pushes every passing change to production automatically.
Basics refresher
CI/CD is continuous integration and continuous delivery or deployment: automation that builds, tests, and ships code on every change. A registry is a versioned store of built artefacts. On-call is the rota of people paged when production breaks.
Three questions decide whether you need an LLM at all
If the input is tabular, use classical ML. If it is language but the task is narrow and high-volume, a small tuned model is cheaper and faster. Reserve LLMs for open-ended work that needs judgement, and prototype with an LLM before deciding to specialise.
Why a Head of Product cares
- LLM cost at scale is the number that surprises finance. A narrow task at a million calls a day almost always wants a small model.
- Prototype with the LLM anyway: it is the fastest way to learn what the task really needs before investing in a specialised model.
- Mixed architectures are normal: an LLM for the conversation, classical models for the scores it consults.
Go deeper
LLMs can act as zero-shot classifiers, which makes them ideal for prototyping and for labelling the data a small model later trains on (LLM-assisted labelling). Small language models (under ~10B parameters) fine-tuned on that data often match the LLM on the narrow task at a fraction of latency and cost (ByteByteGo's LLM vs SLM comparison). Embedding models plus a classifier are another cheap middle path for text. The decision should be revisited as prices fall.
Basics refresher
A small language model is an LLM-style model with far fewer parameters, often run on modest hardware. Zero-shot means the model performs a task from an instruction alone with no examples. Entity extraction pulls names, dates, and amounts out of text.
Embeddings turn meaning into coordinates, which turns search into geometry
An embedding model maps text (or images, or products) to a point in a space with thousands of dimensions, arranged so that similar meanings are close together. Finding relevant documents becomes finding nearest neighbours, and that is the engine under RAG, semantic search, and recommendations.
Why a Head of Product cares
- Embeddings are cheap and reusable: one embedding pipeline powers search, deduplication, clustering, and recommendations.
- Retrieval quality is mostly embedding quality plus chunking. Evaluate retrieval separately from generation.
- Embedding models are swappable, but switching one means re-embedding everything; plan the migration.
Go deeper
Embedding models are trained so that texts with similar meaning have high cosine similarity. Approximate nearest-neighbour indexes (HNSW and similar) make search fast at millions of vectors. Embeddings of different types can share a space (CLIP puts images and captions together), enabling image search by text. Hybrid search combines vectors with keyword scoring so exact terms (product codes, names) still match. Evaluate with recall@k: for known question-passage pairs, how often is the right passage in the top k results?
Basics refresher
A dimension is one axis of a space; the picture has two, real embeddings have hundreds or thousands. Nearest neighbour means the stored point with the smallest distance to the query point. Semantic search finds by meaning rather than by matching words.
Inference cost is a set of knobs, and half of them cost you no quality at all
Smaller, quantised, or distilled models trade a little quality for a lot of speed and cost. Caching, batching, and speculative decoding keep outputs identical and still cut the bill. Know which knob your team turned before you accept a quality regression.
Why a Head of Product cares
- Ask for output-preserving optimisations first; they are pure savings.
- A quality drop from a smaller or quantised model is only real if your eval set says so. Benchmarks are not your task.
- Interactive and batch workloads want different settings; separate them so batch jobs do not slow the chat product.
Go deeper
Prefix caching stores the computed attention state for a shared prompt prefix (your system prompt and few-shot examples) so only the new part is processed. Continuous batching fills the GPU with many sequences at different stages. Speculative decoding lets a small draft model propose several tokens that the large model verifies in one pass; outputs are identical to the large model alone. Flash attention reorganises the attention computation to use GPU memory more efficiently. Quantisation-aware training recovers most of the quality lost by low precision. Whitepaper example: a 2Γ speed-up for a 2% accuracy drop on a detection task, an acceptable trade in most products.
Basics refresher
Precision here means how many bits store each number (32-bit float vs 8-bit integer), not the classification metric. Cache means keeping a computed result to reuse instead of recomputing. Batch means processing many items together.
Glossary
Questions to ask your team
- What exactly is the target variable, and what decision does the prediction feed?
- Where do the labels come from, how many do we have, and what does each new one cost?
- Was the test set separated in time from training? Has anyone tuned against it?
- What is the baseline, and by how much does the proposed model beat it?
- Are we bias-limited or variance-limited? Show me the learning curve before we buy more data.
- Which pipeline stage caps our accuracy? Have we swapped in ground truth for each stage and measured?
- If a human expert's output scores worse on our metric than the model's does, is the metric wrong?
- Show me the confusion matrix. What does each red cell cost us in dollars?
- Who sets the threshold, and when was it last reviewed?
- What is monitored in production, and what triggers a retrain?
- Can we roll back to the previous model version in under an hour?
- For the LLM parts: which knobs have we turned, and did evals change?
Reading list
- Stanford CS229: Machine Learning, Andrew Ng (Autumn 2018) β YouTube playlist, 21 lectures of about 80 minutes each. The canonical university course under this module. For a product leader the three to watch in full are Lecture 1 (framing), Lecture 8 (bias, variance, data splits) and Lecture 12 (debugging and error analysis); the map below points the rest at the section they support. Lectures 13 to 19 (EM, factor analysis, ICA, control theory) are beyond what this module needs.
- ByteByteGo (Alex Xu) β A Crash Course in CI/CD (the top lane of section 10); Good Code vs. Bad Code (why the code around a model matters as much as the model); Understanding Database Types (where vector stores fit); Large vs Small Language Models (section 11's middle branch); How to Shrink a Language Model Without Making it Too Dumb (quantisation and distillation, section 13); How Large Language Models Learn (training, visually).
- Foundational Large Language Models & Text Generation (Google, 2025) β kaggle.com. Sections on training, fine-tuning, and the inference trade-offs behind section 13.
| Module section | CS229 lecture | Take from it |
|---|---|---|
| 01 Three kinds of learning | Lecture 1, Lecture 5, Lecture 16 | Mitchell's T-E-P definition; generative vs discriminative; what an MDP is |
| 03 What training is | Lecture 2, Lecture 3 | Gradient descent on linear and logistic regression, batch vs stochastic |
| 04 Bias vs variance | Lecture 8, Learning theory section | Learning curves, regularisation, why parameters demand data |
| 05 Data is the product | Lecture 8 | Split sizes by data volume, k-fold cross-validation, feature selection |
| 06 The model zoo | Lecture 9, Lecture 6, Lecture 7 | Trees, bagging vs boosting; SVMs and kernels |
| 07 Neural networks | Lecture 10, Lecture 11 | Layers and backprop; initialisation, normalisation, Adam |
| 09 Diagnosing failures | Lecture 12, Lecture 20 | Error analysis, ablation, optimiser-vs-objective; RL diagnostics |
| Unsupervised toolkit (in 01) | Lecture 13, Lecture 15 | k-means and mixtures; PCA |