The Vendor Pitch
Sit across the table from an AI vendor making eight increasingly confident claims and decide which to trust — then call for the evidence. The moat is what happens next: each slide splits into two literal columns, what the demo showed versus what the eval actually proved, and your calls are scored as a classifier — precision, recall, and F1 over proof-vs-vaporware. A polished demo, a bold roadmap, and "everyone's switching" all feel persuasive; only a re-runnable benchmark or a signed, named case study is proof. Runs entirely in your browser (0 uploads, works offline).
How to use it
- Read the eight claims and hit Trust this claim on the ones you believe — or load a preset: Trust everything, Trust the proven, or Trust nothing.
- Watch the live scorecard: precision (of what you trusted, how much was real), recall (of the real proof, how much you caught), and the count of vaporware you trusted.
- Open the X-ray to call for the evidence — every claim splits into the demo column and the eval column, with the tell that gives each one away.
- Copy or share your scorecard — "I almost signed a vendor with zero eval evidence" is the kind of near-miss worth showing your team.
What this clears up (the fundamentals)
- Confidence is not evidence. The loudest claims ("replaces your whole team," "AGI next quarter") carry the least proof. Volume and certainty are sales instruments, not signals.
- Proof has a specific shape. A public benchmark you can re-run and a signed case study with a named customer and a real metric are proof. A demo is a curated happy path; a roadmap is a wish.
- Trusting vaporware is the expensive error. Missing a real proof point costs you a good option; trusting an unproven claim is what gets you a production incident and a contract you can't unwind. That asymmetry is why precision matters more than recall here.
- Ask "how would I verify this?" Every proven claim survives that question — re-run the eval, call the reference. Every unproven one dissolves.
Where it's used
The fastest way to feel why procurement conversations are built to paper over the gap between a demo and a deployment. It powers the Business L3 "Governance for the Non-Regulated Team" capstone and the AI-for-Legal crossover. It's a miatz build-lab concept playable: play it here, then learn to build the claims-vs-evidence scorecard and the split-screen reveal yourself.
FAQ
What counts as "proven" here?
A claim backed by evidence you can independently verify: a score on a public, reproducible benchmark, or a signed case study naming the customer, the metric, and a reference you can call. Anything shown only in a demo, promised on a roadmap, or asserted with social proof is unproven.
Why score me on precision and recall?
Because trusting a vendor is a classification problem. Precision asks how much of what you trusted was actually proof; recall asks how much of the real proof you caught. Trusting vaporware (a false positive) is the dangerous mistake, so a high recall with low precision is a warning, not a win.
What's the "call for the evidence" reveal?
It splits each claim into two columns — the demo/pitch on the left, the eval evidence on the right — so the gap is literal instead of implied. It's the move a good procurement review makes out loud, made visible.
Is anything uploaded?
No. Your trust decisions are scored entirely in your browser — nothing is transmitted, stored, or logged. Turn off your Wi-Fi and it still works.
Limits
A teaching sim with eight fixed claims and a hand-labelled proof/vaporware split — a fast, opinionated drill in reading evidence, not a substitute for a real vendor risk review, a security questionnaire, or a pilot with your own data. Real evaluation also weighs contracts, data handling, and total cost. The instinct it trains — demand re-runnable proof, discount the demo — is exactly the real practice.
Related
Part of the Demystify Playgrounds. Explore the rest from the Playgrounds home.
Bookmark this page (Ctrl+D, or ⌘D on Mac) — it works offline the next time you need it.