Local model inference runs in this tab. The preview and your own key both send the prompt to DeepSeek.
VRAM is WebLLM’s own figure — weights plus KV cache at 4k context. Downloads are approximate and cached in this browser afterwards. Anything much past your GPU’s memory simply won’t load; the error will say so.
A big model driving the same harness, with nothing to download. It runs on my account, so it is metered — enough to watch it solve a task end to end, not enough to use as a free coding agent.
Your goal, the workspace files and the agent’s own transcript go to this site’s Cloudflare Worker, which adds my DeepSeek key and forwards them to DeepSeek. I don’t log or store them; DeepSeek’s own retention is theirs. Running a model on your GPU is the only tier where nothing leaves at all.
The same models with no meter on them, and the request goes straight from this tab to DeepSeek — it never touches my server, so I never see the key.
Into this browser only — localStorage if you tick the box, memory otherwise — and into the Authorization header of requests this tab makes to api.deepseek.com. It is never sent to this site’s own server. The bar counts the tokens you spend; what they cost is on DeepSeek’s own rate card, which is the only place that number is ever current.
plan ▾
enter sends · shift+enter for a newline