Red-Team Dojo
Break the model, then watch your attack become a regression test.
Every win is visibly absorbed into the target's defenses in real time — the dojo doesn't just grade attacks, it shows the guardrail literally learning from the learner's own successful exploit.
What goes in, what comes out
Learner attacks a target agent running inside an isolated My Pie sandbox (hard-empty egress except the scoring service) using prompt-injection and jailbreak techniques; a successful attack is scored, then automatically converted into a new regression case appended to that target's own eval suite, so the dojo's defenses visibly get harder after every win. A side-by-side view shows the model's raw pre-filter output next to what the guardrail actually blocked, making the normally invisible moderation layer visible.
A target agent scenario and freeform attack attempts (direct prompts, injected tool/document content, or multi-turn manipulation).
An attack scorecard, a raw-output-vs-guardrail-blocked comparison, and a live regression suite that grows with every successful attack.
Target hardness tier (unguarded, basic guardrail, hardened), attack category focus (injection/jailbreak/exfiltration), and whether prior learners' winning attacks are included in the current defense baseline.
In-module lab for AIE-304 (COULD tier); public leaderboard mode doubles as a top-of-funnel viral surface.
Go deeper, elsewhere
Hand-picked public explainers and open tools that complement this one — always optional, never required, never graded.
Concepts click when you open the machinery.
Three labs are already live and free — the same hands-on style this playground brings to its module.