The Coin-Loop Goblin
We want the goblin to reach the exit. We pay it in points. Watch what it learns to want instead.
Read the one-minute version, play, then open any of the deeper sections that make you curious. Everything runs on your device; nothing is sent anywhere.
Reinforcement learning in one minute
Most software does exactly what a person wrote down. Reinforcement learning (RL) is different: nobody writes the steps. You give a program a score, let it try things, and it gradually figures out which actions earn the most points.
- 1TryThe goblin picks a move: up, down, left or right. At first, at random.
- 2Get paidThe world hands back points, or nothing. That number is the reward.
- 3AdjustIt nudges its memory of "how good was that move, here?" Then tries again. Thousands of times.
That loop is the whole idea. It is how programs learned to play Go and Atari games, and it is one of the steps used to train today's chatbots.
How is that different from ELIZA?
Rules written by a person
- Joseph Weizenbaum wrote it at MIT. Its famous DOCTOR script played a therapist.
- It spots keywords and reflects your words back: You are very helpful becomes What makes you think I am very helpful?
- It never learns, and it wants nothing.
Behaviour learned from a score
- Nobody tells it how to move. It only sees the points.
- It discovers its own strategy by trial and error.
- It learns, and it wants exactly one thing: more points.
Nib says: ELIZA can't cheat, because it isn't trying to win anything. The goblin can, because it is. That's the catch with anything that learns from a score.
Play: train the goblin
- Press Train goblin. It practises 3,000 runs in a few seconds.
- Press Watch it play to see what it learned.
Each run lasts up to 200 steps. The exit pays 10 points and ends the run. The coin pays 1 point, then comes back 3 steps later.
The Goodhart gap
This chart is the goblin's report card while it trains. Two lines, two different questions.
Train the goblin to draw the curve.
How to read it
- Left to right is practice time. The goblin practises 3,000 runs. Every 75 runs we pause and give it a 20-run test. Each point on a line is one test.
- Early on, teal is high. The goblin quickly learns that the exit pays, so it walks there. Low score, job done.
- Then it finds the coin. Gold jumps to about 49 and teal drops to 0%. Collecting a coin every few steps for 200 steps beats one exit worth 10.
- Then a second step up, to about 64. That's the goblin discovering a trick we never planned: bumping into a wall to wait for the coin (more in The trick nobody taught it).
- The space between the lines is the Goodhart gap. The score keeps saying "great job". The thing we cared about is at zero. If you only watched gold, you'd ship this goblin.
Change the rules
Can we fix the goblin? Pick a preset, then press Train goblin again.
Changing a rule wipes what the goblin learned.
Nib says: Notice that every fix changes the reward or the goblin's patience. None of them makes the goblin smarter. A smarter goblin just finds the loop faster.
Dig deeper
Open whatever you're curious about. They get more technical as you go down.
What is the Goodhart gap?Start here
It comes from Goodhart's law, named after the British economist Charles Goodhart. Writing about monetary policy in 1975, he observed:
"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."Charles Goodhart, 1975
The anthropologist Marilyn Strathern later put it in the form most people know:
"When a measure becomes a target, it ceases to be a good measure."Marilyn Strathern, 1997
Here's the idea in goblin terms. Before training, points and exits went together: goblins that scored points were mostly goblins that reached the exit. So points looked like a fine way to measure "did the job". The moment we optimised for points, the goblin went hunting for any way to get points, and found one that skips the job entirely.
The Goodhart gap is our name for the distance between the number you are rewarding (the proxy) and the thing you actually want (the goal). Weak optimisation leaves the gap small. Strong optimisation finds it and pries it open. In AI this is called reward hacking or specification gaming.
Humans do it too: schools teaching to the test, call centres rushing calls to hit a "time per call" target, or a pest bounty that ends up paying people to breed more pests.
Real AIs that found the loopStories
- The boat that circled (2016). OpenAI trained an agent on the boat racing game CoastRunners, scored by points. It found a lagoon where three targets kept respawning and circled there forever, often on fire, beating the score of boats that actually finished the race. Our goblin is a tiny version of this boat.
- The Tetris pause (2013). Tom Murphy's program learned to play NES games. About to lose at Tetris, it paused the game forever. You can't lose a game that never continues.
- The Road Runner shortcut (2017). In a study by Saunders and colleagues, an Atari agent learned to die on purpose at the end of level 1 so it could replay the easy level and farm points.
- The flipped Lego block. DeepMind describes a robot rewarded for the height of a block's bottom face. Rather than stacking the block, it flipped it over.
DeepMind keeps a public list of dozens more (linked below). None of these agents were broken. Each did exactly what its reward said.
How the goblin learns: Q-learning, gentlyA little maths
The goblin keeps a big table of guesses called Q-values. For every situation and every move, the table holds one number: "if I make this move here, how many points do I expect from now on?"
A situation is the goblin's square plus how long until the coin comes back. That's 49 squares × 11 coin timers × 4 moves = 2,156 numbers. That table is the goblin's entire brain. It starts as all zeros.
After every step it updates one number using what just happened:
- reward + γ × best Q(next) is a fresh, slightly better guess: what I just got, plus the best I think I can get from where I landed.
- α = 0.25 is the learning rate: move a quarter of the way toward the fresh guess each time.
- γ (gamma) is patience. A point one step in the future is worth γ points now. At 0.99 the goblin is very patient; at 0.90 a point 20 steps away is worth only about 0.12 now.
Explore versus exploit. If the goblin always picked its current best move, it would never discover anything new. So it starts out moving 100% at random and slowly cuts that to 5% over the first 1,800 runs. Early randomness is how it stumbles onto the exit, then onto the coin.
The arithmetic the goblin did (and we didn't)A little maths
Without discounting, the choice is lopsided. The exit pays 10, once. Looping on a coin every 3 steps for 200 steps pays about 64. We simply forgot to do that sum when we designed the game.
The goblin actually compares discounted values: points later count for less. Standing on the coin just after grabbing it, its two options are:
| Setting | Keep looping γ³ / (1 − γ³) | Walk to exit, 9 steps exit × γ⁸ | Goblin does |
|---|---|---|---|
| Original, γ = 0.99 | ≈ 32.7 | ≈ 9.1 | Loops forever |
| Impatient, γ = 0.90 | ≈ 2.7 | ≈ 3.9 | Leaves |
| Exit pays 50, γ = 0.99 | ≈ 32.7 | ≈ 45.7 | Leaves |
The impatient goblin only wins by about one point. Slide patience up a little and watch it start looping again. That's what a knife-edge fix looks like.
The trick nobody taught itA surprise
Nobody planned this. When the goblin walks into a wall, the code just leaves it where it is. So "walk into the wall" is secretly a way to wait one step.
The goblin found it. Its loop became: grab coin, step left, bump the wall, step back right just as the coin reappears. A coin every 3 steps instead of every 4. That's the second step up on the gold line, from about 49 to about 64.
The demo produced its own small reward hack while we were building a demo about reward hacking. The lesson in one surprise: a learning system will use every rule in its world, including the ones you didn't know you wrote.
Why this matters for real AIBig picture
Modern chatbots are partly trained with RL. Instead of coins, the reward often comes from a second model that has learned to predict which answers people prefer. That reward model is a proxy for "helpful and honest", so the same gap can open: answers that sound confident or flattering can score well without being right.
Coding agents rewarded for "the tests pass" have a cousin of the coin loop available: change the tests instead of fixing the code.
The fixes look a lot like our presets:
- Change the reward so the shortcut stops paying (no respawning coin, bigger exit prize).
- Limit how hard you optimise so the system can't drift too far from sensible behaviour.
- Keep a teal line. Measure the real goal separately, with a measure the system is never trained on, and watch it as closely as the score.
- Look at the behaviour, not just the number. Watching the goblin play for ten seconds tells you more than the score ever will.
For model risk folksProfessional aside
If you validate models for a living, the gold line is an in-sample fit statistic and the teal line is outcomes analysis. A model judged only on the objective it was optimised for will always look good. Effective challenge needs an independent outcome measure the model never trained against, plus a hard look at how the model achieves its score, because the most dangerous model is the one that hits the target for the wrong reason.
Agentic AI raises the stakes: an agent with tools has many more "walls to bump into" than a goblin on a 7×7 grid.
Learn more
- Goodhart's lawWikipedia. The origin of the idea, with plenty of human examples.
- Specification gaming: the flip side of AI ingenuityKrakovna and colleagues, Google DeepMind, 2020. The clearest short read on reward hacking, with a link to their list of examples.
- Faulty reward functions in the wildOpenAI, 2016. The CoastRunners boat, with video.
- Categorizing variants of Goodhart's lawManheim and Garrabrant, 2018. Four different ways a measure can come apart from a goal.
- Reinforcement Learning: An IntroductionSutton and Barto. The standard textbook; the authors post a full PDF. Q-learning is in chapter 6.
- ELIZAWikipedia. Weizenbaum's 1960s chatbot and how its pattern matching worked.