Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

What a machine actually learns

Machine Learning, Foundations · lesson 1 of 9 · 8 min

Fitting, not understanding

Here are four flats in Lagos, with monthly rent in naira.

| Floor area (m²) | Rent (₦/month) | |---|---| | 40 | 180,000 | | 55 | 240,000 | | 70 | 310,000 | | 90 | 380,000 |

You want a program that guesses the rent of a flat it has never seen. You could write rules by hand. Instead you write down a shape with blanks in it:

rent = w * area + b

w and b are the blanks. They are called parameters. Every choice of w and b is a different program. Learning is the search for the pair that works best.

"Best" has to be a number

Guess w = 4000, b = 0. Run the four flats through it:

| Area | Predicted | Actual | Off by | |---|---|---|---| | 40 | 160,000 | 180,000 | 20,000 | | 55 | 220,000 | 240,000 | 20,000 | | 70 | 280,000 | 310,000 | 30,000 | | 90 | 360,000 | 380,000 | 20,000 |

Average error: ₦22,500. That single number is called the loss. It is the only thing the machine can see about how it is doing.

Now try w = 4000, b = 20000:

| Area | Predicted | Actual | Off by | |---|---|---|---| | 40 | 180,000 | 180,000 | 0 | | 55 | 240,000 | 240,000 | 0 | | 70 | 300,000 | 310,000 | 10,000 | | 90 | 380,000 | 380,000 | 0 |

Average error: ₦2,500.

The loss fell from 22,500 to 2,500. That drop is the entire content of the word "learning". No insight was gained. Two numbers moved and a third number got smaller.

The three ingredients

Every supervised model, from a two-parameter line to a language model with a trillion parameters, is made of the same three things.

  1. 1A shape with parameters. A line. A tree of if-statements. A stack of matrix multiplications. You choose the shape.
  2. 2A loss. One number saying how wrong the current parameters are, averaged over your data. You choose the loss, and the choice matters more than people expect.
  3. 3A procedure for changing the parameters to lower the loss. For a straight line you can solve it with algebra. For anything larger you walk downhill, which is Lesson 8.

That is the whole mechanism. Everything else in this course is about doing it honestly.

What the model does not have

The fitted model has no concept of a building. Hand it area = 500 and it will report ₦2,020,000 without hesitation, because it has never been told that flats of that size are rare, or that the market above 200 m² behaves differently. It has never been told anything. It has four rows and two knobs.

This also means the shape you picked is a hard ceiling. Suppose rents flatten out above 120 m², because very large flats sell rather than rent. A straight line cannot bend. No quantity of extra data will make w * area + b produce a curve — more data will only pin down the best possible straight line more precisely, and the best possible straight line is still wrong. If you want a curve, you have to pick a shape that can curve.

This is why "which model should I use" is a real question and not a detail. You are not asking which program is smarter. You are choosing which family of functions the search is allowed to look inside.

Parameters and hyperparameters

w and b are fitted from the data. But other numbers get chosen before fitting starts: how many terms the polynomial has, how deep a tree may go, how large each downhill step is. These are hyperparameters. They are picked by you, by trial, and where you are allowed to run those trials is the subject of Lesson 4.

Keep the distinction clear from the start. Parameters come from the data. Hyperparameters come from you.

Before you move on

You fit `rent = w * area + b` to 50,000 flats. The true relationship curves: rent rises steeply up to 120 m² and then almost flattens. Training finishes. What has it accomplished?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly

What a machine actually learns · Machine Learning, Foundations · Addaly