Kinds of prediction — study guide
The concept's fragments, read in order.
What are you actually predicting
The first question to ask about a machine learning problem is not which model to reach for but a plainer one: what is the shape of the thing you are trying to predict? Machine learning already assumes you will learn a rule from labeled examples instead of writing it by hand, and the families of problems sort themselves by the kind of answer you want out the other end.
There are only a few of these families, and they nest. The top split is whether your examples come with known answers at all. When they do, the question becomes whether the answer is one of a fixed set of categories or a number on a sliding scale. When they do not, the work becomes finding structure that no one labeled in advance.
Naming the family is the cheap step that saves the expensive ones. It decides what data you must collect, what a correct answer even looks like, and how you will know the model is any good — so it is worth getting right before you touch anything else.
With an answer key, or without
The widest division in machine learning is whether your examples carry answers. Supervised learning works from labeled data: every example comes with the known answer you want the model to produce, so training can check each prediction against the truth and correct itself toward it. Given photos already marked cat or dog, the model learns the mapping from image to label.
Unsupervised learning has no such key like studying with an answer key you can check each guess against, versus being handed a pile of unlabeled material and asked to find the patterns yourself. The data is a pile of examples with features but no labels attached, and the goal is to surface structure that is already latent in it — which examples resemble each other, how the data spreads out — without anyone having said in advance what the right grouping is.
The distinction is not a stylistic preference; it follows from what you have. If you can get labels, supervised methods can learn to reproduce them. If labels are missing, too expensive, or the point is to discover categories nobody has drawn yet, the unsupervised side is where you work.
Which category does it belong to
Classification is the supervised task whose answer is a category: given an example, predict which of a fixed, finite set of classes it belongs to. Is this email spam or not, is this image a cat or a dog, which of ten digits is this — each is a choice among labels decided ahead of time like sorting each item into one of a fixed set of labeled bins. The output is discrete, one label out of a known list, never a value in between.
The simplest case is binary, a choice between two classes. When the list runs longer — ten digits, a hundred species — the task is multiclass, but the shape is the same: still a pick from a fixed set. What makes something a classification problem is that the set of possible answers is enumerable and settled before training begins.
That discreteness is the whole signature. If you can write down the complete list of answers and each example gets exactly one of them, you are classifying.
How much, as a number
Regression is the supervised task whose answer is a continuous number rather than a category. The price of a house, tomorrow's temperature, the steering angle a car should hold — none of these is a pick from a fixed list; each is a value on a scale, and any point on that scale is a legal answer like reading a value off a continuous dial rather than choosing from labeled bins.
That is the one line that separates it from classification. Classification chooses among a finite set of labels; regression reads out a quantity, and between any two answers there is always another. The model is not sorting an example into a bin but estimating how much.
The difference is not academic bookkeeping. A continuous target lets the model express that one prediction was almost right and another was wildly off by how far the number missed — a notion of nearness that a bag of unrelated category labels simply does not have.
Grouping without labels
Clustering is an unsupervised task: it groups similar examples together without ever being told what the groups are. The data arrives as features with no labels, and the algorithm's job is to find examples that resemble one another and gather them, so the categories emerge from the data rather than being handed to it.
Because nobody supplied the answer, there is nothing to check a grouping against the way supervised training checks a prediction. The structure clustering finds is proposed, not confirmed — a claim that these examples belong together because they sit close in the space of their features, which may or may not match a division you care about.
That is exactly why it earns its place. When you suspect there are natural groups in your data but no one has drawn them yet, clustering is how you get a first look at the structure before you have any labels at all.
Beyond the basics
Classification, regression, and clustering cover a great deal of ground, but they are not the whole map. Other families exist, each defined by a different kind of output or a different kind of feedback, and it is worth knowing their names even before you know their machinery.
Ranking orders a set of items by relevance rather than labeling each one. Recommendation picks, from a large catalog, the items to put in front of a particular user. Generation produces new content as its output instead of a label or a number. Reinforcement learning learns from rewards and penalties earned by acting in an environment rather than from a fixed set of labeled examples.
Each of those is a subject in its own right, and naming them here is deliberately all this does. The point is only that the first three families are a starting map, not the border of the territory — the shape of the output keeps defining new tasks well past where this concept stops.
The same data, framed differently
The same dataset does not come with a task stamped on it. A table of house records can be posed as regression — predict the sale price as a number — or as classification — predict whether a house sells above or below some threshold. The data is identical; the question you decide to ask is what makes it one task or the other.
That choice is not free, because the framing sets everything that follows. It fixes what labels you must produce, since a category and a continuous value are different targets to collect. It fixes the loss the model minimizes, the score that says how wrong a prediction is. And it fixes how you measure success, because you judge a category guess and a numeric estimate by different standards.
So the task is a decision you make, not a property you discover. Pose the same data as a different task and you have changed the labels, the objective, and the yardstick all at once — which is why it belongs at the very front of the work.
Framing decides the build
Consider the steering of a self-driving car. Frame it as classification and the model chooses among a few labels — turn left, go straight, turn right — a discrete verdict for each moment. Frame it as regression and the model outputs a precise steering angle, a continuous number that can land anywhere on the range the wheels allow.
These are not two flavors of the same build; they are different builds. They need differently labeled data, a coarse set of turn categories versus recorded angles down to the degree. They optimize different losses and are judged by different measures. And they behave differently on the road, since a handful of turn buckets can only approximate the smooth correction an angle expresses directly.
That is why the task family is the load-bearing first decision. Choose it and you have already shaped the data you must gather, the model you can train, and the standard it will be held to — long before a single parameter is fit.