How often does pure noise look significant?
Every world in this grid is boring by construction — a random A/B split with zero real effect. Run the same significance test on all 100 and watch a few light up anyway. That feeling is what a p-value actually promises.
Every world is an A/B split drawn from the exact same 8% conversion process — zero real effect, by construction. The same two-proportion test runs on all of them.
Watching a handful of random, meaningless worlds turn green out of 100 is the fastest way to feel — not just be told — what a p-value actually promises: a false-alarm budget, never a truth detector.
What does a p-value actually tell you?
How likely data at least this extreme would be if nothing real were going on. It is not the probability your effect is real — that's the single most common misreading in dashboards and papers alike.
What is a null world?
A dataset generated with genuinely no effect — both groups drawn from the exact same process. Any 'significant' result inside one is a false alarm by construction, which is what makes them the perfect intuition drill.
Why do some effect-free worlds still light up as significant?
That's the deal you sign at p<0.05: roughly a 5-in-100 false-alarm budget. Run 100 tests on pure noise and about five clear the bar. Run enough experiments in a quarter and noise alone will hand you a 'winner.'
This playable is part of the registry: its tool page
Next: will that model fit on your GPU?
Can You Actually Run This? shows the VRAM bill line by line — weights, KV cache and overhead against real GPU memory sizes.