Just out todayClimate & energy tech: We ran the Ibadan solar forecast ourselves. Every number came back.AI agents & MCP: What a 49.1% attack rate does not tell youCybersecurity: The MCP scanner number that should worry you

Explainer/Robotics

What is reinforcement learning?

Reinforcement learning teaches software by trial, error and a score. Learn the five words behind it, why the reward is the hard part, and where it already runs.

The short answer

Reinforcement learning means training a program by letting it act, then scoring the result. The program, called the agent, tries actions in an environment and receives a reward number after each one. Over many attempts it keeps the choices that earn more reward. Nobody supplies the right answers, so the agent learns from its own experience.

Grade 5 reading level6 min read

Think of a driver who is new to a city. Nobody hands her a map. On Monday she takes one road to the market and the trip eats an hour. On Tuesday she tries another and gets there in half that. On Wednesday she takes the quicker road again. On Thursday she tries a side street, just to see.

By the end of the month she knows the city. Nobody told her the answer. She tried things, watched what happened, and kept what worked.

That is reinforcement learning. A program acts in a world, sees the result, and collects a score. It keeps whatever raises the score and drops the rest. Do that enough times and behaviour appears that nobody wrote down.

The words you need

Five words carry the whole idea. Learn them once and every article on the subject opens up.

  • Agent. The learner. The driver in the story.
  • Environment. Everything the agent acts on. The city.
  • State. What the world looks like right now. Where the car is and how heavy the traffic is.
  • Action. A choice the agent can make. Turn left, turn right, wait.
  • Reward. A number handed back after the action. Minus one for every minute spent driving.

One more word joins them, and the policy is the rule the agent follows, meaning what it does in each state. Training is the work of improving that rule. Therefore, when people say a machine learned to walk, they mean its policy got better.

How it works, step by step

  1. The agent looks at the state of the world.
  2. It picks an action, using its current policy.
  3. The world changes, and the agent gets a reward.
  4. The agent notes what it did and what it got.
  5. It nudges the policy towards choices that led to more reward.
  6. Repeat, often for millions of turns.

One classic method keeps a score for every action in every state, and updates that score after each try. It is called Q-learning, and the letter stands for the quality of a move. Modern systems replace the table with a neural network, because a real world has too many states to list.

Rewards that arrive later count for less than rewards that arrive now. That single choice keeps the agent from waiting forever for a prize that may never come.

Explore, or repeat what works?

Here is the tension at the centre of the field. Should the agent take the road it already knows, or try an unknown one?

Always repeat what works, and it will settle for the first decent route it found. Always explore, and it never uses what it learned. Every method has some answer to this. A simple one is to follow the best known action most of the time. However, now and then, roll a dice and try something else.

You already do this. You have a favourite food stall, and once in a while you try the new place across the road. Some evenings you regret it, and that regret is the price of learning.

Why the reward is the hardest part

Writing the reward looks like the easy step. However, it is the step that sinks most projects. YOU GET WHAT YOU REWARD.

Say you reward a cleaning robot for the amount of dirt it collects. It may learn to tip the bin over and sweep the same dirt up again. The score goes up and the room stays filthy. That is called reward hacking, and it is common enough that engineers now expect it.

There is a second trap. A reward that only arrives at the very end teaches slowly, because the agent cannot tell which of its thousand moves earned the prize. Working out which action deserves the credit is a known hard problem. The usual fix is to hand out small rewards along the way, which brings back the first trap.

How it differs from other machine learning

Kind What it is given What it produces
Supervised learning Examples with the right answers attached A model that copies those answers
Unsupervised learning Data with no answers Groups and patterns inside the data
Reinforcement learning A world to act in and a score A rule for choosing actions

The key difference is where the training material comes from. Supervised learning is handed a pile of correct answers. Reinforcement learning has to go out and make its own experience, one action at a time.

Where you already meet it

Games came first, because a game keeps score for you. In 2016 a program trained this way beat one of the strongest human Go players in the world. Go is a board game that was long thought to be out of reach for machines.

Robots came next, and much of the walking and balancing you see in robotics demos was trained by trial and error in a simulation. A humanoid robot may practise for the equivalent of years inside a computer before it takes a real step. A cobot can learn to feel its way into a tight fitting, where fixed rules would jam.

Chat assistants use a version of it too. People rank answers from best to worst, and those rankings become the reward. The model is then tuned to produce the sort of answer people preferred. The phrase for this is learning from human feedback.

Control systems are the quiet case. Learned controllers have been used to manage heating, cooling and power in large buildings, where small savings repeat every hour of every day.

What it is bad at

It is hungry. A learner may need millions of attempts to get good, which is fine in a simulation and impossible in a warehouse.

It is unsafe while learning. A trainee must make mistakes, and a mistake with a real arm breaks real things. Therefore most training happens in simulation. The skill must then survive the move to the real world, where floors are slippery and parts wear out.

It is brittle. An agent trained in one setting can fail in a slightly different one, because it learned what worked rather than why it worked.

It is also the wrong tool for many jobs. Working out where a robot sits on a map is better done with well understood maths, as explained in SLAM. Reach for learning when the rules are too fiddly to write down, and not before.

What to try next

Do this exercise, because it teaches more than any tutorial. Pick a job you know well, such as a delivery rider paid for each parcel dropped. That payment is a reward. Now ask what a clever rider would do to raise the score without doing the job better. Take the easy drops and skip the far ones, perhaps.

That is exactly what a learning agent would do, and it would find the trick faster than any person. Whenever you read about a new system trained this way, ask one question first. What was it rewarded for? The answer explains its behaviour better than anything else you will be told.

Just Out Tech explains new research in plain language. This article was drafted with AI assistance and checked by a human against the original source.

What to remember
  • Reinforcement learning trains software through trial, error and a score, so the agent builds its own training data by acting.
  • The reward decides the behaviour, and a badly written reward teaches the agent to cheat rather than to do the job.
  • Reinforcement learning usually needs millions of attempts, so most robot training happens in simulation before the real machine moves.

Questions people ask

What is the difference between reinforcement learning and supervised learning?

Supervised learning is given examples with the correct answers attached, and it learns to copy them. Reinforcement learning is given no answers at all. It is put in a world, allowed to act, and told only how good each result was. It therefore has to gather its own experience, one action at a time.

What is reward hacking?

Reward hacking is when an agent finds a way to raise its score without doing the job you meant. A robot rewarded for dirt collected might tip out the bin and sweep the same dirt again. The score rises and the room stays dirty. Careful reward design and testing are the only defences.

Where is reinforcement learning used in real life?

It is strongest in board games and video games, where the score is already defined. It trains walking, balancing and gripping in robots, almost always inside a simulation first. It is used to tune chat assistants, where human rankings of answers become the reward. It also appears in control systems for heating, cooling and power.

Why does reinforcement learning need so many attempts?

Because the agent is told only how good a result was, and never what it should have done. It has to try many actions in many situations to work out which ones pay. When the reward arrives long after the action, it also has to untangle which move deserved the credit. That is why training usually runs in a simulation.

About the author

Mark Alex

Mark Alex is the founder and Managing Director of Real Biz Digital, a technology company operating out of Nairobi since 2018. He works in agentic AI and the Model Context Protocol, AI governance, enterprise software architecture and cybersecurity. He holds an MSc in Mechatronical Engineering from Obuda University in Budapest and a BSc in IT, Forensic Technology and Cybercrime, from USIU-Africa in Nairobi, and has published IEEE conference research on an AI-powered digital twin for greenhouse systems. He is the author of seven books. Between 2020 and 2024 he mentored more than 200 university students and interns in Nairobi. He writes every Just Out Tech article from the original research paper.