Miatz playground

Train the Judge

You train the judge, then we expose its bias.

Coming to the Build-LabAIE-401BYOK AIAI-Concept-Vizdev · data
The Demystify signature

Makes the learner the annotator, then shows their own aggregated preferences producing a biased judge -- judge drift stops being abstract and becomes personal and undeniable.

How it works

What goes in, what comes out

What it does

Learner sees ten pairs of AI responses to the same prompt and picks the 'better' one each time, exactly as a human preference-labeler would. The game then trains a simple simulated reward model on those choices and shows it scoring new, trickier responses, surfacing a reward-hacking example where a fluent-but-wrong answer scores high because of a bias the learner's own picks introduced.

You bring

Ten pairwise 'this one's better' clicks.

You get

A personal reward-model bias readout, e.g. 'you rewarded confident tone 80% of the time, even when it was wrong', and a live reward-hacking demonstration.

You control

Prompt-domain picker (helpfulness, safety, code correctness); optional sabotage mode that secretly favors length/confidence to demonstrate how bias creeps in unnoticed.

Where it's used

Concept companion inside AIE-401 (LLM-as-judge / rubric design); lead magnet.

RLHFpreference pairsreward modelingjudge drift/bias

Concepts click when you open the machinery.

Three labs are already live and free — the same hands-on style this playground brings to its module.