Can You Actually Run This?
Pick a model and your GPU. We do the honest math.
the formula is shown working, not just the verdict — every term in parameter-count times bytes-per-parameter plus KV-cache is visible and adjustable, replacing 'just download and run it' hype with arithmetic anyone can rerun themselves
What goes in, what comes out
Learner picks a model's parameter count, a quantization level, and a GPU (dropdown of common consumer/prosumer cards or a custom VRAM entry). The tool runs the real arithmetic live — params times bytes-per-parameter, plus KV-cache-per-token times context length, plus runtime overhead — and renders a fuel-gauge VRAM bar with a pass/fail verdict and an estimated tokens/sec range.
three clicks — model, quantization, GPU — then drag the context-length slider and watch the bar fill live
a literal fits/doesn't-fit verdict with every number in the formula shown, an estimated tokens/sec range, a suggested quantization if it doesn't fit
model-size selector, quantization toggle (fp16/int8/int4), context-length slider, batch-size slider, GPU picker or custom VRAM input
FREE, top-of-funnel — /playgrounds gallery; feeds directly into AIE-110's graded VRAM-arithmetic exercise
Go deeper, elsewhere
Hand-picked public explainers and open tools that complement this one — always optional, never required, never graded.
Concepts click when you open the machinery.
Three labs are already live and free — the same hands-on style this playground brings to its module.