Skip to content
VellumOpen Vellum

Ask your pages. Nothing leaves the device.

Vellum’s assistant is off until you turn it on. Then a language model downloads once and runs in your browser. It answers from what you wrote, shows where it found it, and can run code for you, with your approval.

Vellum’s assistant panel, explaining that the model downloads once and runs on this device, with the button to turn it on.
Vellum’s assistant panel, explaining that the model downloads once and runs on this device, with the button to turn it on.

On your device, after you opt in

Nothing AI-related downloads until you turn the assistant on. Then its model is fetched once from its public host and cached by the browser; it runs on the GPU with WebGPU, or as a smaller model on the CPU. Your questions, the answers and the search index it builds stay in the browser’s storage on this device.

The model is a small open-weights model running locally: it is slower and less capable than large cloud models, and it can be wrong. Its citations show you where to check.

Answers with sources

Every answer cites the pages it used as numbered chips; a chip opens the page. Ask across all pages, one notebook or only the open page. Files count too: text inside PDFs, Word documents and slides is indexed page by page, so a citation can point at the exact page of a PDF.

More than chat

In the app

Questions

Does the assistant send my pages anywhere?

No. It runs in your browser. The only network use is downloading the model once, from its public host, after you turn it on.

What does it need?

A current browser. It is fastest with WebGPU (recent Chrome and Edge); without it, a smaller model runs on the CPU.

Can it change my pages?

Only when you ask for a summary, outline or continuation and keep it. Code it wants to run waits for your approval.