How to play
Each machine pays out a coin with its own hidden probability. You have 100 pulls. Pull any machine you like; with Show statistics on, each card shows its observed payout rate and a 95% interval for the true rate. At the end you see the true rates and how the algorithm played.
Same luck, different strategy. Every machine has a pre-drawn sequence of outcomes. Your third pull of machine B gives exactly the same result as the algorithm’s third pull of machine B, so the only difference between your score and its score is the choices you made.
The explore–exploit dilemma
Every pull on a machine you are unsure about is a pull not spent on the machine that looks best. Explore too little and you may commit to a worse machine; explore too much and you waste pulls on machines you already know are worse. This is the same trade-off faced by A/B tests, clinical trials, ad systems and recommendation engines.
How the algorithm thinks
The opponent uses Thompson sampling. For each machine it keeps a belief — a beta distribution built from its wins and losses so far. Before each pull it draws one plausible payout rate from every belief and pulls the machine with the highest draw. Machines it knows little about have wide beliefs, so they occasionally draw high and get tried; machines that are clearly worse stop being chosen.
Can you beat it? What the simulations say
We pitted the algorithm against a simple human-like strategy — try each machine twice, then always pull the best-looking one — over thousands of games with identical luck:
| Easy (3 machines) | Hard (5 close machines) | |
|---|---|---|
| Coins expected from always pulling the best machine | 66.9 | 59.1 |
| Thompson sampling, average | 60.2 | 51.5 |
| “Try twice, then commit”, average | 60.3 | 53.0 |
| Random pulling, average | 48.5 | 47.6 |
| Ends up on the best machine: Thompson vs commit | 93% vs 77% | 52% vs 45% |
Over a short 100-pull game, committing early does as well as the algorithm on average — but it locks onto the wrong machine far more often, so its bad games are worse. Over a longer horizon the algorithm pulls clearly ahead: in 1,000-pull games on the hard level it lost about 36 coins to exploration against 48 for the commit strategy. So yes, you can beat it — but you have to judge when you have explored enough.
Connections
The interval bars on the cards are Wilson confidence intervals. Bandits are the adaptive cousin of the A/B test; if you want to monitor an ordinary A/B test as it runs, use a sequential test.
Frequently asked questions
A problem where you repeatedly choose between options with unknown payoffs, such as slot machines, ads or treatments, and want to maximise your total reward. Every choice is a trade-off between exploring options you know little about and exploiting the one that currently looks best.
A Bayesian strategy: for each option it keeps a probability distribution of how good the option might be, draws one plausible value from each, and picks the highest. Uncertain options sometimes draw high values and get tried; options that are clearly worse fade out.
In a short game, yes. Over 100 pulls, exploring briefly and then committing does about as well as Thompson sampling on average, but it locks onto the wrong machine more often. Over 1,000 pulls Thompson sampling clearly wins.
An A/B test explores two versions equally before deciding; a bandit shifts traffic towards the better version while the test runs. Bandits earn more during the experiment, while classic A/B tests give cleaner estimates of the difference.