Model Arena
Read five rounds of anonymized answers — four per round, labeled A to D — and bet which model wrote the best one before any identity is shown. The arena scores each blind pick against the hidden-truth quality, then reveals the real names and the true leaderboard you just calibrated yourself against. The moat is that reversal: instead of handing you a ranking as received wisdom, it makes you generate your own blind comparison first, so the leaderboard becomes something you tested, not something you trusted. Runs entirely in your browser (0 uploads, works offline).
How to use it
- Read each round's anonymized answers (A–D) and judge which one is strongest.
- Click the model you bet wrote that best answer — one pick per round, change it as often as you like.
- Watch your correct / total score and accuracy update live as you commit picks.
- Open the X-ray panel to see your blind card round by round, then copy or share your scorecard.
What this clears up (the fundamentals)
- A leaderboard is a claim, not a fact — every ranking rests on a scoring method. Score the rounds yourself and the ranking stops being an authority you defer to and becomes a measurement you can check.
- Blind evaluation removes the halo — strip the model names and a "big-lab" answer wins no points it didn't earn on the text. Your gut for which brand is "best" is exactly the bias blind tests exist to expose.
- Wins and average quality are different axes — a model can top the most rounds yet not have the highest mean, and vice-versa. The leaderboard shows both so "best" is a question with two honest answers.
- Calibration is a skill you can measure — "7 of 10 blind" is a number, and it moves as you practice reading outputs instead of trusting logos.
Where it's used
A top-of-funnel model-literacy check — the fastest way to feel why a leaderboard is a methodology, not a scoreboard handed down from on high. It's a miatz build-lab concept playable: play the blind round here, then learn to build the scoring rubric, the hidden-quality model, and the leaderboard reveal yourself.
FAQ
What am I actually guessing?
For each round you pick which model, from the fixed roster, you think produced the best answer. The names are hidden while you read, so you're betting on the text — and on your own priors about each model — not on the label.
How is my score calculated?
Each round has one answer with the highest hidden quality; that model is the round's winner. Your pick is correct when it matches that winner. Accuracy is simply your correct picks divided by the number of rounds, rounded — a deterministic rubric, no model call and no randomness.
Why can a model with fewer wins rank above one with more average quality?
The leaderboard sorts by wins first (rounds where a model had the top answer), then breaks ties by average quality across all rounds. A model that wins two rounds outranks one that wins one, even if the single-win model is stronger on average — which is exactly the kind of nuance a one-number ranking hides.
Is anything uploaded?
No. Every answer, its hidden quality, your picks, and the scoring all live in your browser — nothing is transmitted, stored, or logged. Turn off your Wi-Fi and it still works.
Are these real model outputs?
No — the roster and answers are a fixed teaching set with authored quality scores, so the exercise is deterministic and reproducible. The method (blind read, commit a pick, reveal and score) is exactly how real head-to-head evaluations like LMArena-style testing work.
Limits
A teaching arena with five fixed rounds, a small roster, and hand-set quality — it trains the habit of blind evaluation, not a verdict on any real model. Genuine leaderboards aggregate thousands of votes or graded tasks with confidence intervals, and quality is multi-dimensional (accuracy, latency, cost, safety). The move it teaches — score blind before you read the ranking — is the real practice.
Related
Part of the Demystify Playgrounds. Explore the rest from the Playgrounds home.
Bookmark this page (Ctrl+D, or ⌘D on Mac) — it works offline the next time you need it.