Miatz playground

Model Arena

Blind-guess which model answered, then see the real leaderboard.

Coming to the Build-LabAIE-130BYOK AIAI-Coredev · data
The Demystify signature

The market sells hype and rankings as received wisdom; this makes the learner generate and score their own blind comparison first, so the leaderboard becomes something they've personally calibrated against, not just trusted.

How it works

What goes in, what comes out

What it does

Learner submits their own prompt and gets back anonymized responses labeled A/B/C/D from different model families; before any reveal, they guess which label is which model and rate the outputs, building a personal prediction history. On reveal, real identities plus live leaderboard data (Artificial-Analysis/LMArena-style Elo, price, context window) populate next to the learner's guess, and every hover on a jargon term (MoE, distillation, quantization) links to a one-line, sourced definition.

You bring

A learner-authored prompt, and the learner's blind guesses/ratings for each anonymized response.

You get

A reveal screen mapping labels to real models, a running personal calibration score (e.g. 7/10 correctly identified), and live leaderboard stats per model with jargon tooltips.

You control

Number of models in the blind round (2-6), model pool (closed + open-weight mix), and whether leaderboard data refreshes live or uses a cached daily snapshot.

Where it's used

Free top-of-funnel lead magnet (no signup needed for one round) and in-module lab for AIE-130 (elective, Model Literate badge).

model landscape literacyjargon (MoE, distillation, context window, Elo)blind evaluationleaderboard methodologycalibration tracking

Concepts click when you open the machinery.

Three labs are already live and free — the same hands-on style this playground brings to its module.