AI Engineering · 400-level
Evals & Quality Systems
AIE-4013 creditscorebadge: evaluatorprereqs: AIE-203, AIE-301
Start this course — free
Earn the evaluator credential
What's inside
Sections & lessons
01
Golden datasets & rubric design
- Writing the 20 cases a system must passconcept45 min
02
Code graders vs. LLM-as-judge
- Deterministic checks next to a judge model, run togetherconcept60 min
03
The disagreement heat-map
- Where the code grader and the AI judge part waysconcept60 min
04
Grader-hacking
- Crafting an output that fools the judge but still fails the code graderconcept60 min
05
Reward-model bias & judge drift
- Teaching a simple reward model your own bias, then seeing itconcept60 min
06
Quiz: eval-design judgment
- Checkpoint on rubric gaps and grader-hacking riskconcept30 min
07
Lab + eval-gate: golden set + grading suite
- 20-case golden set, two graders, threshold pass (Evaluator badge)concept120 min
Learn it from the inside
This module's playgrounds
Vocabulary
Key concepts in this course
Compare related approaches
Optional · watch & try
Go deeper, elsewhere
Hand-picked public explainers and open tools — always optional, never required, never graded.
Introduction to LLM Evaluations – Model Evals vs Application EvalsDistinguishes model-level evals from application-level evals Intro to LLM Evaluation w/ OpenAI Evals [Walk-Thru]Hands-on walkthrough building a golden-set eval How to Systematically Setup LLM Evals (Metrics, Unit Tests, LLM-as-a-Judge)Covers code graders vs. LLM-as-judge tradeoffs PromptfooFree open-source tool for running golden-set evals and red-team tests against prompts
This module ends in a gate you can fail.
That's what makes passing it mean something. Take the DSAT, get placed, and start earning.