Run
Reward
Press Train
Exploring
Exploration
How to play
- Pick Maze or Cart-pole.
- Press Train. The robot practises over and over, fast. Watch the reward chart climb.
- In the maze, tap squares to add walls and pits. Can it still find the way?
- Press Watch to see it practise in real time, one move at a time.
| Do this | Phone | Keyboard |
|---|---|---|
| Train fast / stop | Train | T |
| Watch real time / stop | Watch | W |
| Pause / go on | Train or Watch | P |
| Forget everything | Reset | R |
| Wall, pit, empty | Tap a square (again to change) | Arrows, then Enter (on the maze) |
| Poke the pole | Tap beside it, or ← Poke / Poke → | ← → (on the stage) |
| Scene | Maze / Cart-pole | 1 2 |
| Speed | Slider | [ ] |
What's happening?
- Nobody tells the robot what to do. It only gets a reward: +20 for the goal, -20 for a pit, -1 for every move. On the cart: +1 for every step the pole stays up, -100 when it falls.
- It keeps a big table of guesses, called Q: "if I'm here and do this, how much reward will I get in the end?" At first every guess is 0.
- After each move it nudges one guess towards reward now + γ × the best guess from where it landed. That's Q-learning. The learning rate α is how big the nudge is.
- The discount γ says how much future reward counts. Near 1: plan far ahead. Near 0: only care about right now.
- Exploration ε is the chance it tries a random move instead of its best guess. It needs some, or it never finds out there's something better. Fading it lets it explore early and show off later.
- In the maze, brighter squares are worth more and the arrows are its best guess. Good news spreads back from the goal one square at a time.
- The cart-pole has too many exact states for a table, so each number (where, how fast, angle, spin) is sorted into a few bins. 864 boxes in all.
- This is reinforcement learning, the same idea that taught computers to play Go and robots to walk. Big ones swap the table for a neural network, like the one in the Neural Net Playground, trained by sliding downhill as in Gradient Descent Hill.
Try this: pick Cart-pole, set Exploration to 0 and press Reset, then Train. It only ever does what it already thinks is best, so it gets stuck. Now slide Exploration back up and Reset. How long can it balance now?