Explainer/Artificial intelligence
What is machine learning?
Machine learning is how a program picks up a skill from examples instead of from rules. Here is the training loop in plain words, the three main kinds, and the checks to run before you trust a model.
Machine learning means a computer program that gets better at a task by being shown examples, rather than by being given rules. You supply data with the right answers attached. The program adjusts its own internal numbers until it gets most of the examples right. It then applies what it found to new cases it has never seen.
A child learns what a goat is without ever hearing a definition. You point at goats. You point at sheep, and you say the word each time. After a hundred animals the child is right nearly every time. However, ask that child for the rule and no rule comes out.
Machine learning is the same thing done by a computer. You show a program many examples with the right answer attached. The program tunes itself until it gets most of them right, and then it handles animals it has never seen.
This is the engine inside nearly all modern artificial intelligence. Learn it once and most AI news stops being mysterious.
What machine learning actually is
Ordinary programming runs one way, and a person thinks up the rules and writes them down. Data goes in and answers come out.
Machine learning runs the other way. Data goes in together with the answers, and the rules come out. Those rules are called a model.
Stop on the word model. A model is not a document you can read. This is because it is a long list of numbers, often millions of them, that turns an input into an output. Nobody picked those numbers by hand, and training picked them.
How training works, step by step
Every training run follows the same loop.
- Split your examples into two piles. A big pile to learn from and a small pile to test with.
- Start the model with random numbers, so its first guesses are useless.
- Show it one batch of examples, and compare each guess with the right answer.
- Measure how wrong it was, and that measure is called the loss.
- Nudge every number a little in the direction that lowers the loss.
- Repeat, often millions of times, until the loss stops falling.
Then comes the test that matters, so run the model on the small pile it never saw. Why keep a pile back? Because a model can memorize its lesson without learning the pattern. Scoring it on the examples it studied tells you nothing at all.
The three main kinds
Almost every project fits one of three shapes.
Supervised learning uses examples with the answer attached. Photos marked cat or dog, loans marked repaid or not repaid. This is the common kind, and it needs the most human work up front.
Unsupervised learning uses examples with no answers, and the program groups things that look alike. A shop can use it to find that its buyers fall into four natural groups, without deciding the groups in advance.
Reinforcement learning uses reward in place of answers. The program tries something, gets a score, and tries again. This is how programs learn games and how robots learn to walk.
Why the data matters more than the method
New teams ask which algorithm is best. However, that is the wrong first question. Ask what is in the data.
A plain method on good, plentiful, well labeled data will beat a clever method on thin, messy data almost every time. This is why teams spend most of their weeks gathering, cleaning and labeling. The training data sets the ceiling, and the method only decides how close you get to it.
What it is good at
Machine learning fits a narrow shape of problem, and three things must be true at once.
- The pattern is really there in the data, even if no person can write it down.
- You have many past cases, and you know the right answer for them.
- Tomorrow will look roughly like yesterday.
When all three hold, the results can be very good. Reading handwriting, spotting fraud, guessing which customer is about to leave, sorting fruit on a belt by size and color.
Where it goes wrong
The classic failure has a name, and overfitting means the model learned the noise along with the signal. It scores beautifully on its lesson and badly in the field, and a student who memorizes past exam papers hits the same wall.
The second failure is drift, because the world moves and the model does not. A fraud model trained before a new scam appeared will wave that scam through. Therefore, retrain on fresh data on a schedule, and watch the score over time.
The third failure is the quiet one. The model works, and it works for a reason you did not intend. Say every photo of a sick lung in your set came from one hospital machine. Therefore, the model may end up reading the machine rather than the lung.
Where you already meet it
Machine learning is already in your day. Your email sorts spam with it, and your bank flags odd payments with it. Your phone opens with your face because of it. Mobile money systems use it to catch accounts that move stolen cash. Shops use it to guess how much stock to hold next week.
Deep learning is the branch that gets the headlines. It uses a neural network with many layers, and it is what made image, speech and text work well. A large language model is a very big example of it.
How to start, and what to check
Do not start with a model. Start with a decision somebody makes by hand every week, and with the record of what they decided. That record is your data set.
Then ask four questions before any money is spent. Do we hold at least a few thousand past cases? Do we know the right answer for them? Could a person do this job from the same information? What does a wrong answer cost us?
If a wrong answer is expensive, keep a person in the loop and let the model advise. THE DECISION COMES FIRST AND THE MODEL SECOND. Get that order right and the rest is craft.
Just Out Tech explains new research in plain language. This article was drafted with AI assistance and checked by a human against the original source.
- Machine learning is the method behind almost every working AI system today, and it swaps hand written rules for patterns found in data.
- A machine learning model must be judged on data it has never seen, because scoring it on its own training examples proves nothing.
- Most of the effort in a machine learning project goes into gathering and cleaning data rather than into choosing the algorithm.
Questions people ask
What is the difference between machine learning and artificial intelligence?
Artificial intelligence is the broad goal of getting machines to do work that needs human thought. Machine learning is one method for reaching it, and it is the method behind nearly all working systems today. So every machine learning system is AI, and not every idea labeled AI uses machine learning.
How much data do I need for machine learning?
There is no single number, and it depends on how hard the pattern is. A simple task with clear signals may work with a few thousand labeled examples. Image and language tasks usually need far more. A good rule is to start with the data you already hold, measure the score, and see whether more data moves it.
What is overfitting?
Overfitting is when a model learns the quirks of its training examples instead of the general pattern. It scores very well on data it studied and poorly on anything new. You catch it by holding back a portion of your data and testing on that. You reduce it with more examples, a simpler model, or both.
Can a machine learning model explain its decisions?
Only in part. Simple models such as decision trees can be read directly. Large models cannot, because the decision is spread across millions of numbers. Tools exist that show which inputs mattered most for one case. They give a useful hint rather than a full reason, so keep a person in the loop for decisions that carry weight.