Hub

Gradient Descent Hill

Step 0
Rate 0.1
Loss (height) over steps log scale

How to play

  1. Pick a landscape: Bowl, Valley, Bumpy or Saddle.
  2. Tap anywhere on the map to drop three balls there. They race to the bottom.
  3. Slide the learning rate up until something blows up.
Do thisPhoneKeyboard
Drop the ballsTap or drag the mapArrows, then Enter (on the map)
Play / PausePlayP
One stepStepN
Race againRestartR
Learning rateSlider[ ]
LandscapeBowl / Valley / Bumpy / Saddle1 to 4
Racer on / offTap its cardTab to it, Enter

What's happening?

  • The map is a landscape seen from above. Dark = low, light = high. Each line joins spots of the same height.
  • A ball can't see the whole map. It only feels the slope right under it, then takes a step downhill. Then again.
  • The learning rate is the size of each step. Too small: it crawls. Too big: it jumps over the bottom, and can bounce out for ever.
  • Plain GD just steps downhill. Momentum is a heavy ball that keeps its speed. Adam takes about the same size step in every direction.
  • Training an AI is the same game, in millions of directions at once. The height is the loss: how wrong the AI is.
  • Want to see it for real? The Neural Net Playground runs this exact trick to teach a tiny network.

Try this: pick Bumpy. Can you find a learning rate where Plain GD gets stuck in a little dip, but Momentum rolls through to the real bottom in the middle?