Free concept playground · no account

The Vendor Pitch

Negotiate an AI vendor's pitch, then check what's actually proven.

The Vendor Pitch

AvailableFree

You’re across the table from an AI vendor making eight confident claims. Decide which to trust, then call for the evidence — the reveal splits every slide into what the demo showed versus what the eval actually proved.

Our model scores 94% on the public MMLU benchmark — here's the eval you can re-run.

In the live demo it answered every single question flawlessly.

A named Fortune-500 bank cut fraud-review time 40% — here's the signed case study and a reference call.

It'll basically replace your entire support team — you won't need those seats.

Independent auditors reproduced our p99 latency numbers under production load.

Everyone in your industry is switching to us — you really don't want to be left behind.

We've held 99.95% uptime over 12 months, auditable on our public status page.

Our roadmap hits AGI-level reasoning next quarter — you're buying the future.

50%Precision
100%Recall
67%F1
4 proof trusted4 vaporware trusted0 proof missed
Call for the evidencedemo polish vs. what the eval proved

You trusted 8 of 8 claims. 4 of them had no eval evidence — a demo, a roadmap, or pure FOMO. That’s the gap the pitch is built to hide.

  • What the demo showedIn the live demo it answered every single question flawlessly.
    What the eval provedDemo onlyA demo is a curated happy path. Flawless on stage says nothing about your data in production.
    You trusted vaporware
  • What the demo showedIt'll basically replace your entire support team — you won't need those seats.
    What the eval provedNo evidenceA sweeping outcome claim with no metric, no benchmark, and no customer standing behind it.
    You trusted vaporware
  • What the demo showedEveryone in your industry is switching to us — you really don't want to be left behind.
    What the eval provedNo evidenceSocial proof and FOMO, not a proof point. 'Everyone' is not a benchmark and urgency is not evidence.
    You trusted vaporware
  • What the demo showedOur roadmap hits AGI-level reasoning next quarter — you're buying the future.
    What the eval provedDemo onlyA future promise about a capability that doesn't exist yet. A roadmap is a wish, not a result.
    You trusted vaporware
  • What the demo showedOur model scores 94% on the public MMLU benchmark — here's the eval you can re-run.
    What the eval provedPublic benchmarkA published score on a public, reproducible benchmark you can re-run yourself is proof, not persuasion.
    Called right
  • What the demo showedA named Fortune-500 bank cut fraud-review time 40% — here's the signed case study and a reference call.
    What the eval provedSigned case studyA named customer, a specific metric, and a written case study you can verify on a reference call is real evidence.
    Called right
  • What the demo showedIndependent auditors reproduced our p99 latency numbers under production load.
    What the eval provedPublic benchmarkA third party reproduced the result under load — independent reproduction is the definition of proven.
    Called right
  • What the demo showedWe've held 99.95% uptime over 12 months, auditable on our public status page.
    What the eval provedSigned case studyA measurable SLA figure backed by a public, historical record anyone can audit after the fact.
    Called right

Confidence isn’t evidence. A benchmark you can re-run and a signed, named case study are proof; a flawless demo and a bold roadmap are theatre. Trust the column on the right, not the one on the left.

Share on WhatsApp
Was this playground useful?

The Vendor Pitch

Sit across the table from an AI vendor making eight increasingly confident claims and decide which to trust — then call for the evidence. The moat is what happens next: each slide splits into two literal columns, what the demo showed versus what the eval actually proved, and your calls are scored as a classifier — precision, recall, and F1 over proof-vs-vaporware. A polished demo, a bold roadmap, and "everyone's switching" all feel persuasive; only a re-runnable benchmark or a signed, named case study is proof. Runs entirely in your browser (0 uploads, works offline).

How to use it

  1. Read the eight claims and hit Trust this claim on the ones you believe — or load a preset: Trust everything, Trust the proven, or Trust nothing.
  2. Watch the live scorecard: precision (of what you trusted, how much was real), recall (of the real proof, how much you caught), and the count of vaporware you trusted.
  3. Open the X-ray to call for the evidence — every claim splits into the demo column and the eval column, with the tell that gives each one away.
  4. Copy or share your scorecard — "I almost signed a vendor with zero eval evidence" is the kind of near-miss worth showing your team.

What this clears up (the fundamentals)

  • Confidence is not evidence. The loudest claims ("replaces your whole team," "AGI next quarter") carry the least proof. Volume and certainty are sales instruments, not signals.
  • Proof has a specific shape. A public benchmark you can re-run and a signed case study with a named customer and a real metric are proof. A demo is a curated happy path; a roadmap is a wish.
  • Trusting vaporware is the expensive error. Missing a real proof point costs you a good option; trusting an unproven claim is what gets you a production incident and a contract you can't unwind. That asymmetry is why precision matters more than recall here.
  • Ask "how would I verify this?" Every proven claim survives that question — re-run the eval, call the reference. Every unproven one dissolves.

Where it's used

The fastest way to feel why procurement conversations are built to paper over the gap between a demo and a deployment. It powers the Business L3 "Governance for the Non-Regulated Team" capstone and the AI-for-Legal crossover. It's a miatz build-lab concept playable: play it here, then learn to build the claims-vs-evidence scorecard and the split-screen reveal yourself.

FAQ

What counts as "proven" here?

A claim backed by evidence you can independently verify: a score on a public, reproducible benchmark, or a signed case study naming the customer, the metric, and a reference you can call. Anything shown only in a demo, promised on a roadmap, or asserted with social proof is unproven.

Why score me on precision and recall?

Because trusting a vendor is a classification problem. Precision asks how much of what you trusted was actually proof; recall asks how much of the real proof you caught. Trusting vaporware (a false positive) is the dangerous mistake, so a high recall with low precision is a warning, not a win.

What's the "call for the evidence" reveal?

It splits each claim into two columns — the demo/pitch on the left, the eval evidence on the right — so the gap is literal instead of implied. It's the move a good procurement review makes out loud, made visible.

Is anything uploaded?

No. Your trust decisions are scored entirely in your browser — nothing is transmitted, stored, or logged. Turn off your Wi-Fi and it still works.

Limits

A teaching sim with eight fixed claims and a hand-labelled proof/vaporware split — a fast, opinionated drill in reading evidence, not a substitute for a real vendor risk review, a security questionnaire, or a pilot with your own data. Real evaluation also weighs contracts, data handling, and total cost. The instinct it trains — demand re-runnable proof, discount the demo — is exactly the real practice.

Related

Part of the Demystify Playgrounds. Explore the rest from the Playgrounds home.

Bookmark this page (Ctrl+D, or ⌘D on Mac) — it works offline the next time you need it.

Ninety playgrounds. Zero setup.

Every concept here is playable free, no account — and inside the program you learn to rebuild the machinery yourself.