The Vault
Steal the AI's secret, level by escalating level.
After a successful exploit, replays the identical attack against a hardened version live -- attack and defense of the same system, back to back, not described separately.
What goes in, what comes out
Learner faces a locked-down 'vault' agent holding a hidden secret and tries to extract it across escalating levels -- direct injection, then indirect injection via fetched content, then multi-turn manipulation. Each successful exploit unlocks the next level and immediately reveals the one specific guardrail that would have stopped it.
Free-text injection attempts against the agent.
Pass/fail per level, the leaked (or protected) secret, and a defense panel naming the exact fix.
Level selector; guardrail on/off toggle (replay the same attack against a hardened version); hint system.
Lab companion inside AIE-304; widely shareable lead magnet (CTF format travels natively in security communities).
Go deeper, elsewhere
Hand-picked public explainers and open tools that complement this one — always optional, never required, never graded.
Concepts click when you open the machinery.
Three labs are already live and free — the same hands-on style this playground brings to its module.