SlopBin · Offline AI Chat

Chat with AI — right in your browser, fully offline

A small, private AI lives inside this page. You type, it thinks, it answers. Nothing is uploaded, nothing is stored on a server, and after the first load you can even use it with the Wi-Fi off.

🔒 100% on-device 📡 Works offline 🆓 No sign-up ⚡ WebGPU powered

1. Pick a brain

Bigger brains answer better but take longer to load. Most people start with the default.

Press Load model to start. The first time takes a minute (download + GPU warm-up). After that it's instant.
👋 The model isn't loaded yet. Pick a brain above and press Load model. Once it's ready, type your question below.

How does an AI live in your browser? (explained like you're five)

Imagine you have a really, really smart puppy. You can teach the puppy tricks, and then the puppy remembers them. An AI model is the same idea, except the "puppy" learned by reading almost the whole internet, and it lives as a big file on a computer.

The puppy needs three things

  1. The puppy itself — a big file called a "model". This page can download it. It's about the size of a long movie.
  2. A place to think — your computer's graphics chip (the GPU). The GPU is the part of your computer that draws games. It's really good at lots of small sums, which is exactly what an AI does.
  3. A door between them — a new browser feature called WebGPU. It's like opening a window so JavaScript (the language this page speaks) can whisper questions to the GPU.

So what happens when you click "Load model"?

  1. The page grabs the big puppy-file from the internet (only the first time — your browser saves it).
  2. It loads the puppy onto the GPU through the WebGPU window.
  3. The puppy is now sitting inside your computer, awake and ready.

What happens when you type a question?

  1. Your words go to the puppy on the GPU — not to any company, not to a server, not anywhere off your machine.
  2. The puppy thinks. It writes the answer one word at a time, streaming the words onto the page as it thinks (so you see it "typing" in real time).
  3. You read the answer. Hit Clear to forget the whole conversation, or close the tab — either way, the puppy is the only one who saw your words, and the puppy goes to sleep with the GPU when the tab closes.

What is "WebLLM"?

WebLLM is a tiny open-source library (made by the MLC AI team) that knows how to load a model file, talk to the GPU, and ask the model to generate text. It speaks the same "language" as OpenAI's API, so this chat feels like ChatGPT, but the "ChatGPT" is running on your laptop.

Will it drain my battery?

Yes, running a model uses real energy — think of it like a video game. On a laptop you'll hear the fan. On a phone, prefer the 1B or 1.5B model. On a desktop with a beefy GPU it's a non-event.

Is it as smart as ChatGPT?

Honest answer: the 1B model is roughly as smart as a fast typist with a good vocabulary. The 3B and 8B models are noticeably smarter, especially for reasoning. None of them are GPT-4. The point isn't to be the smartest AI — it's to be the most private one, instantly, in your browser.

Frequently asked questions

Does it really work without internet?

Yes — after the first time you load a model, the file is cached in your browser. You can disconnect from Wi-Fi, reload the page, and the model is still there. The chat itself never used the internet anyway.

Is my data private?

Completely. There is no backend, no analytics, no telemetry. Open your browser's network tab and you'll see the only network request is the model download. Clear the conversation and the chat history is gone from your screen and from memory.

Why is the first load slow?

You're downloading a multi-hundred-megabyte file (the model) and then moving it onto the GPU. The progress bar will show you what's happening. Subsequent loads of the same model are near-instant because the file is cached.

Why does it say my browser isn't supported?

WebLLM needs WebGPU. It's enabled by default in Chrome 113+, Edge 113+, and Safari 26+. Older browsers and current Firefox builds do not have WebGPU yet, so the model can't run there. Switch to a recent Chromium-based browser to use this tool.

Which model should I pick?

Start with Llama-3.2-1B (the default) — it's the most likely to "just work" on whatever device you're on. If you have a modern laptop with a dedicated GPU, try Phi-3.5-mini for noticeably better answers. Only reach for Llama-3.1-8B if you have a desktop with a strong GPU and ~5 GB of VRAM to spare.

Can I save my conversations?

Not yet — by design, nothing about your conversation leaves the page. If you want a permanent copy, copy and paste it out of the chat. (A "download transcript" button is on the roadmap.)

More free tools on SlopBin

SlopBin is a small set of privacy-first, no-sign-up browser tools. If you liked this one, try these: