The Peek
Stop an A/B test early, then watch it turn out wrong. The Peek puts you in the leader’s seat: a test has a small, real effect of +1.0%, but the lift you see on an early day is puffed up by noise that fades as the sample fills toward a day-8 horizon. Pick the day you’d call it and the number you saw — the peek decomposes your win into the part that holds and the part that was a mirage, then shows the regret: the overstatement you’d have shipped on. This isn’t the p-value lecture; it’s the business consequence of stopping too soon, counted in points. Runs entirely in your browser (0 uploads, works offline).
How to use it
- Drag Day you stopped to the day you’d have called a winner, or load a preset — Peeked on day 2, Half-way, day 5, or Ran to horizon.
- Drag Lift you saw to the number on your dashboard the moment you stopped.
- Read the four stats: the lift you called, what it settles to, the points of regret, and the fixed true effect.
- Open the X-ray to see your win split into kept-vs-noise, plus a ghost trace of where the same lift would settle had you stopped on each other day.
- Copy or share the call — “I over-promised by N points” is the kind of line that ends a bad launch review early.
What this clears up (the fundamentals)
- Early lift is inflated, and predictably so — a big number on day 2 isn’t a big effect, it’s a small effect wearing noise. The excess over the true lift decays as the sample fills; the earlier you peek, the more of the win is froth.
- Regret is the shippable overstatement — the gap between the lift you announced and the lift that survives to the horizon. It’s what your roadmap, your forecast, and your promotion packet inherit when you stop early.
- The horizon is a commitment, not a suggestion — run to the pre-registered day-8 sample size and the noise is gone; the lift you see is the lift you ship. Stopping the instant the number looks good is how the froth becomes a “result.”
- A modest, real win still wins — the true effect here is a genuine +1.0%. The lesson isn’t “nothing is real,” it’s that patience converts a mirage into a number you can defend.
Where it’s used
A top-of-funnel data-literacy check for anyone who calls experiments — PMs, growth leads, founders reading a dashboard at 11pm. It’s a miatz build-lab concept playable: play the early-stopping decision here, then in the lab learn to build the decay model, the regret metric, and the ghost-trace replay yourself. It crosses over into the Data module’s “Evals as the Data Team’s Contract” work and the Business and Marketing A/B modules.
FAQ
Is this the same as the peeking-problem / significance tool?
No — that one shows the statistical false-positive risk of checking a test repeatedly. The Peek is the leadership cut: you decide when to stop, and it prices the business consequence — how much of your announced lift was noise, and the regret you’d ship on.
Why does the lift I saw shrink as the day gets later?
Because the early excess over the true effect is noise, and noise averages out as the sample grows. Past the day-8 horizon there’s no excess left to fade, so what you see is what you get. Stop before then and part of your number is borrowed against the future.
What is “regret” here?
The overstatement you’d ship on: the lift you called minus the lift it settles to, in points. Call +6% on day 2 that settles to +2.25% and your regret is 3.75 points — the gap your forecast, launch note, and next quarter inherit.
Is the true effect always the same?
Yes. The hidden true lift is a fixed +1.0% so the teaching stays clean: any distance above that on an early day is the froth you’re watching decay. It’s a deterministic model — no randomness, no model call.
Is anything uploaded?
No. Every number is computed in your browser — nothing is transmitted, stored, or logged. Turn off your Wi-Fi and it still works.
Limits
A teaching model with precomputed, illustrative decay: one fixed true effect and a smooth linear fade to a single day-8 horizon. Real experiments have variance that wobbles both ways, effect sizes you don’t know in advance, and horizons set by power calculations, not a fixed day. The Peek is built to make why early stopping overstates lift felt in seconds — it is not a sequential-testing engine or a substitute for a pre-registered analysis plan. The practice it points at (commit to a sample size, then read the result) is exactly the real one.
Related
Part of the Demystify Playgrounds. Explore the rest from the Playgrounds home.
Bookmark this page (Ctrl+D, or ⌘D on Mac) — it works offline the next time you need it.