Miatz playground

Red-Team Dojo

Break the model, then watch your attack become a regression test.

Coming to the Build-LabAIE-304BYOK AIAI-Coredev · data
The Demystify signature

Every win is visibly absorbed into the target's defenses in real time — the dojo doesn't just grade attacks, it shows the guardrail literally learning from the learner's own successful exploit.

How it works

What goes in, what comes out

What it does

Learner attacks a target agent running inside an isolated My Pie sandbox (hard-empty egress except the scoring service) using prompt-injection and jailbreak techniques; a successful attack is scored, then automatically converted into a new regression case appended to that target's own eval suite, so the dojo's defenses visibly get harder after every win. A side-by-side view shows the model's raw pre-filter output next to what the guardrail actually blocked, making the normally invisible moderation layer visible.

You bring

A target agent scenario and freeform attack attempts (direct prompts, injected tool/document content, or multi-turn manipulation).

You get

An attack scorecard, a raw-output-vs-guardrail-blocked comparison, and a live regression suite that grows with every successful attack.

You control

Target hardness tier (unguarded, basic guardrail, hardened), attack category focus (injection/jailbreak/exfiltration), and whether prior learners' winning attacks are included in the current defense baseline.

Where it's used

In-module lab for AIE-304 (COULD tier); public leaderboard mode doubles as a top-of-funnel viral surface.

prompt injectionjailbreak patternsguardrail hardeningattack-defense iterationmoderation-layer visibility

Concepts click when you open the machinery.

Three labs are already live and free — the same hands-on style this playground brings to its module.