Free playground · no account

How often does pure noise look significant?

Every world in this grid is boring by construction — a random A/B split with zero real effect. Run the same significance test on all 100 and watch a few light up anyway. That feeling is what a p-value actually promises.

Every world is an A/B split drawn from the exact same 8% conversion process — zero real effect, by construction. The same two-proportion test runs on all of them.

Pick a threshold and sample size, then run the batch — the grid fills in live as each empty world gets tested.
The Demystify reveal

Watching a handful of random, meaningless worlds turn green out of 100 is the fastest way to feel — not just be told — what a p-value actually promises: a false-alarm budget, never a truth detector.

What does a p-value actually tell you?

How likely data at least this extreme would be if nothing real were going on. It is not the probability your effect is real — that's the single most common misreading in dashboards and papers alike.

What is a null world?

A dataset generated with genuinely no effect — both groups drawn from the exact same process. Any 'significant' result inside one is a false alarm by construction, which is what makes them the perfect intuition drill.

Why do some effect-free worlds still light up as significant?

That's the deal you sign at p<0.05: roughly a 5-in-100 false-alarm budget. Run 100 tests on pure noise and about five clear the bar. Run enough experiments in a quarter and noise alone will hand you a 'winner.'

This playable is part of the registry: its tool page

Next: will that model fit on your GPU?

Can You Actually Run This? shows the VRAM bill line by line — weights, KV cache and overhead against real GPU memory sizes.