Peek at Your Own Risk
Peek early, win often, be wrong nearly half the time.
the engine reruns the same null-true experiment 100 times in the background and shows 'stop when significant' producing false wins roughly a third of the time — optional stopping's invisible math rendered as a public scoreboard
What goes in, what comes out
Learner launches a simulated A/B test on a fake conversion metric where the true underlying lift is secretly zero. A live p-value chart updates as fake traffic streams in, and the learner can 'peek' and stop the moment p<0.05 or wait for the pre-registered sample size; a 20-tests-at-once mode shows the winner's-curse effect across a whole team.
start or stop the test at any point; choose peek-and-stop vs wait-for-sample-size; run the 20-simultaneous-tests mode
a p-value-over-time chart, a scoreboard of how many peeked tests falsely called 'significant', and a false-positive-rate readout across repeated runs
true-lift slider (default 0% for the null-is-true lesson); traffic speed; peeking allowed on/off; run 1 vs 20 simultaneous tests
Module 'The PM's Eval' L3 demo + free lightning-lesson teaser on the diagnostic funnel
Go deeper, elsewhere
Hand-picked public explainers and open tools that complement this one — always optional, never required, never graded.
Concepts click when you open the machinery.
Three labs are already live and free — the same hands-on style this playground brings to its module.