Learning from data — study guide
The concept's fragments, read in order.
Learn the rule, do not write it
The usual way to make a computer do something is to write the rules yourself. A function takes an input, and the code you wrote decides what comes out. That works right up until the rule is one nobody can spell out: tell a photo of a cat from a photo of a dog, and there is no tidy list of instructions that gets it right every time.
Machine learning flips the arrangement. Instead of a person writing the rules, you hand the machine examples and let it derive the rule from them. A model is the mathematical relationship it pulls out of the data, and that relationship is what maps new inputs to outputs. The programmer's job shifts from stating the answer to gathering the examples the answer can be learned from like learning to recognize something from seeing many examples of it, rather than from a written checklist of its features.
That is the whole move, and everything else in this concept follows from it. Once the rule comes from the data rather than from you, the data is what you have to get right.
Rows of examples, columns of features
The raw material a model learns from is a dataset: a collection of examples. It helps to picture it as a table, one you would recognize from any data format. Each row is a single example, and each column describes it. The columns that describe the inputs are the features, and when a dataset carries the known answer for each example, that answer sits in its own column called the label.
A house-price dataset makes the split concrete. Square footage, number of bedrooms, and neighborhood are features, one per column; the price each house actually sold for is the label. One row is one house: its features and its answer, side by side. Give the model enough rows like that and it has something to learn from.
That neat table almost never arrives neat. A dataset is somebody's data-wrangling output before it ever reaches a model, and what the model learns is only as trustworthy as the work that produced those rows.
Fitting the parameters to the examples
A model is not a fixed rule; it is a rule with adjustable numbers inside it. Those numbers are its parameters, and they are what decide how a given input turns into an output. Change the parameters and you change the mapping the model performs.
Training is the process of setting those parameters so the model's outputs match the examples. The machine starts with parameters that are essentially guesses, checks how far its predictions fall from the known answers in the data, and nudges the numbers toward values that fit better. Learning, in this concept, is exactly that search: finding parameters that make the input-to-output mapping agree with the examples.
The details of how the nudging is done are their own subject. What matters here is the shape of it: a model learns by adjusting its own parameters to fit the data it is shown, rather than by being told the rule outright.
Learning from answered examples
The most common way to train a model is to hand it examples that already carry the right answer. Each example has its features and its label, and the label is the answer you want the model to produce for inputs like it. Learning from labeled examples this way is called supervised learning: the labels supervise the model, correcting it toward the answers a human already knows.
The mechanism is a comparison. The model looks at an example's features, produces its own guess, and that guess is measured against the label, the answer known to be correct. A gap between guess and answer is the signal that the parameters need adjusting. Across many labeled examples the model discovers the connections between features and answers that produce correct predictions.
The labels are what make the learning possible, and they are also what make it expensive. Someone had to know, or decide, the correct answer for every example before the model could learn from a single one.
The point is what transfers
A model is never trained for the sake of the examples it trains on; those answers are already known. The point is the new case, the input the model has never seen, where the answer is not written down anywhere. So the real goal of training is to capture the underlying pattern in the data, the regularity that still holds for inputs outside the training set. A model that does this generalizes.
The tempting failure is to memorize instead. A model with enough freedom can fit the training examples perfectly, matching every answer it was shown, and still be useless, because it learned the particular examples rather than the pattern beneath them like studying worked examples to learn the method that solves new problems, rather than memorizing the answers to one test. Memorizing the training data fails the moment a new input arrives, which is the only moment that matters.
So fitting the examples well is necessary but not the target. What you are actually after is the part of what the model learned that transfers to data it has never seen.
Garbage in, garbage out
Because a model learns its rule from the data and nothing else, the data sets the ceiling on how good the model can be. Errors in the examples teach errors. Gaps in the examples become blind spots. Bias in the examples becomes bias in the model, faithfully reproduced. This is the old principle stated for machine learning: garbage in, garbage out like a dish that can only be as good as its ingredients, since no technique rescues spoiled inputs.
The trouble is that the model has no way to know what its data left out or got wrong. It cannot tell a correct label from a mistaken one, or a representative sample from a skewed one; it simply fits whatever it is given. If a dataset only ever shows one kind of case, the model learns that the one kind is the whole world, and it will be confidently wrong about everything the data never mentioned.
No amount of clever training rescues bad data. The most consequential decisions in a machine learning project are often made long before training starts, in what gets collected and what gets left out.
Why it takes so many examples
Machine learning has a reputation for being data-hungry, and the reason follows from what the model is trying to do. From a handful of examples, almost any pattern can be made to fit, including coincidences that happen to be true of those few cases and nothing else. The model has no way to tell a real regularity from a lucky accident when it has seen too little.
More examples, and more varied ones, starve out the accidents. A coincidence that held across a few cases rarely survives thousands of them, so the patterns that persist across a large, varied dataset are the ones more likely to be real. Variety matters as much as volume: examples that cover the range of situations the model will face keep it from mistaking a narrow slice for the whole.
This is why serious machine learning leans so hard on gathering data. The quantity and spread of the examples are often what decides whether the model found the pattern or just the noise.
Your data becomes the behavior
Once you understand that a model's rule comes from its data, a great deal of otherwise mysterious behavior stops being mysterious. A system built on machine learning does not do what someone wrote down; it does what the examples taught it. Its strengths, its blind spots, and its biases are all downstream of the data it learned from.
This is what makes the idea load-bearing rather than academic. A self-driving RC car that only ever trained on well-lit, empty tracks will not know what to do with a shadow or an obstacle, because those cases were never in its examples. An agent pipeline assembled from learned components inherits the habits of whatever data shaped each one. In both, the behavior you observe is a reflection of the data that went in.
So the practical lever is not the algorithm but the examples. The data you collect, and the cases you remember to include, are the behavior you are going to get.