Train the Judge
You train the judge, then we expose its bias.
Makes the learner the annotator, then shows their own aggregated preferences producing a biased judge -- judge drift stops being abstract and becomes personal and undeniable.
What goes in, what comes out
Learner sees ten pairs of AI responses to the same prompt and picks the 'better' one each time, exactly as a human preference-labeler would. The game then trains a simple simulated reward model on those choices and shows it scoring new, trickier responses, surfacing a reward-hacking example where a fluent-but-wrong answer scores high because of a bias the learner's own picks introduced.
Ten pairwise 'this one's better' clicks.
A personal reward-model bias readout, e.g. 'you rewarded confident tone 80% of the time, even when it was wrong', and a live reward-hacking demonstration.
Prompt-domain picker (helpfulness, safety, code correctness); optional sabotage mode that secretly favors length/confidence to demonstrate how bias creeps in unnoticed.
Concept companion inside AIE-401 (LLM-as-judge / rubric design); lead magnet.
Go deeper, elsewhere
Hand-picked public explainers and open tools that complement this one — always optional, never required, never graded.
Concepts click when you open the machinery.
Three labs are already live and free — the same hands-on style this playground brings to its module.