4 min read

WebLLM: Run Language Models Inside the Browser (GitHub, Scanned)

An engine that runs open language models entirely in the browser on WebGPU, with an OpenAI-style API.

WebLLM logo
✅
Scan: safe. Nothing to warn about: a compact TypeScript library with no meaningful hits, no telemetry, and known advisories only in development tooling. Scanned Oct 3, 2026; the full report is below.

WebLLM runs a language model inside a web page. The model downloads once into the browser's cache and then runs on the visitor's own GPU through WebGPU, with no server doing inference and no prompt leaving the device. For developers it looks like the OpenAI API: create an engine with a model name, then call engine.chat.completions.create with messages, with streaming, JSON mode and seeding supported.

It is on this list because it removes the server from private AI features. A website, a Chrome extension or an offline web app can offer chat, summarization or structured extraction at zero inference cost to the publisher. It supports Llama, Phi, Gemma, Mistral and Qwen models compiled for the MLC runtime, runs in web workers or service workers so the page stays responsive, and can be imported from npm or straight from a CDN.

WebLLM comes from the MLC AI project, which also makes MLC LLM, and has about 19,000 stars. It is Apache-2.0 licensed, and a paper describing it is on arXiv. WebLLM Chat, a ready-made chat app built on it, runs at chat.webllm.ai.

  • Repository: github.com/mlc-ai/web-llm
  • Licence: Apache-2.0 (Apache License 2.0)
  • Language: TypeScript. Stars: 19.2K. Forks: 1,388. Last push: Oct 3, 2026.
  • Scan: safe, Oct 3, 2026, commit 7f0f0d1

Who it is for

Web developers who want AI features without paying for inference or handling user data, extension authors, and anyone curious to try a local model with nothing to install.

Getting started

1. Try it with no code in a WebGPU browser such as recent Chrome or Edge

open https://chat.webllm.ai

2. Add the package to your web project

npm install @mlc-ai/web-llm

3. Load a model and chat with it, OpenAI style

import { CreateMLCEngine } from "@mlc-ai/web-llm";
const engine = await CreateMLCEngine("Llama-3.1-8B-Instruct-q4f32_1-MLC");
const reply = await engine.chat.completions.create({ messages: [{ role: "user", content: "Hello!" }] });
console.log(reply.choices[0].message.content);

The first load downloads the whole model, several gigabytes for an 8B model, into browser storage, and it needs a GPU with enough memory. Smaller models such as Qwen2 0.5B or 1.5B load much faster and suit phones and older laptops.

Safety scan

We cloned mlc-ai/web-llm at commit 7f0f0d1 on Oct 3, 2026 and ran the checks described on the GitHub Tools page: credential patterns, decode-and-execute code, install-time scripts, committed binaries, risky CI workflows, every host the code talks to, known vulnerabilities in pinned dependencies, and project hygiene. A person read every hit. This is what we found.

  • No secrets, no bare-IP URLs and no committed binaries across 326 files and about 45,000 lines, 37,000 of them TypeScript. The one pattern hit is a long line in tests/function_calling.test.ts, a test fixture. The only npm hook installs Husky for contributors.
  • Network hosts are Hugging Face for model weights (over 200 references), GitHub for compiled model libraries, and MLC's own sites; we found no analytics. Prompts and replies stay in the browser.
  • 9 known advisories, none critical (7 high), in package-lock.json (475 packages): js-yaml, brace-expansion and braces, used by linting, testing and build tools rather than by the engine at run time.
  • 5 workflows, none using pull_request_target; the one third-party action is not pinned to a commit. Security policy, CodeQL, licence and contributing guide present; no Dependabot.

What the scanner counted

CheckResult
SecretsNone found.
Suspicious code1 pattern hit found and read; every one is listed under the raw findings.
Install-time code1 npm lifecycle script
Committed binariesNone.
CI workflows5 workflows. None use pull_request_target. 1 of 1 third-party action pinned to a tag rather than a commit.
Network hosts26 distinct hosts referenced from source; most often huggingface.co, github.com, platform.openai.com, models.example. No URLs to bare IP addresses.
Known vulnerabilities9 advisories across 465 pinned packages: 0 critical, 7 high, 2 moderate, 0 low. docs/requirements.txt: 7 packages, 0 advisories; package-lock.json: 475 packages, 9 advisories.
Project hygieneHas security policy, CodeQL, licence file, contributing guide. Missing automated dependency updates.
OpenSSF ScorecardNot scored: the project is not in Scorecard's weekly index.

The raw findings

Every hit the scanner wrote out, with a link to the exact line at the scanned commit. Secrets candidates are redacted.

Pattern hits (1)
WhereRuleMatch
tests/function_calling.test.ts:411very-long-line3654 chars
npm lifecycle scripts (1)
  • package.json prepare: husky
Worst known vulnerabilities (9 of 9)
AdvisorySeverityPackageSummary
GHSA-2883-xcg3-v3hhhighjs-yaml@3.15.1js-yaml: maxTotalMergeKeys does not limit CPU use for empty merge sources
GHSA-6j4f-fj2g-mc7phighbrace-expansion@5.0.9brace-expansion: DoS via uncontrolled recursion in parseCommaParts causing stack exhaustion
GHSA-qhr7-859c-m2p7highbrace-expansion@5.0.9brace-expansion: DoS via uncontrolled recursion on nested brace groups causing stack exhaustion
GHSA-6j4f-fj2g-mc7phighbrace-expansion@1.1.18brace-expansion: DoS via uncontrolled recursion in parseCommaParts causing stack exhaustion
GHSA-qhr7-859c-m2p7highbrace-expansion@1.1.18brace-expansion: DoS via uncontrolled recursion on nested brace groups causing stack exhaustion
GHSA-vfj7-8cjw-p6xmhighbraces@3.0.3braces vulnerable to stack-exhaustion denial of service through deeply nested patterns
GHSA-2883-xcg3-v3hhhighjs-yaml@4.3.1js-yaml: maxTotalMergeKeys does not limit CPU use for empty merge sources
GHSA-q2hr-2g5m-vwhrmoderatebrace-expansion@5.0.9brace-expansion: Quadratic-time expansion of the `{a},b}` rewrite causes CPU denial of service
GHSA-q2hr-2g5m-vwhrmoderatebrace-expansion@1.1.18brace-expansion: Quadratic-time expansion of the `{a},b}` rewrite causes CPU denial of service

By the numbers

Stars19.2K
Forks1,388
Contributors56
Commits453
Open issues135
Open pull requests17
Releases5
Latest releasev0.2.85
LicenceApache-2.0
Main languageTypeScript
Project age3 years
Last pushOct 3, 2026
Tracked files326
Lines of code45.4K
Checkout size13 MB

Lines by language: TypeScript 37.4K, JavaScript 2,576, Markdown 1,414, CSS 1,368, HTML 978, JSON 867.

Questions

Is WebLLM free?

Yes. WebLLM is Apache-2.0 and free, and because inference happens on each user's device, there is no per-request cost for you or them. The models it runs are free to download, each under its own licence, such as Meta's Llama licence or Apache-2.0 for Qwen.

Which browsers support WebLLM?

Any browser with WebGPU enabled. Recent desktop Chrome and Edge have it on by default, and support in Safari and Firefox has been arriving in recent versions. If WebGPU is missing, the engine reports an error rather than falling back to the CPU.

Does WebLLM send my prompts to a server?

No. Model weights are downloaded from Hugging Face or the configured host, but prompts and replies are processed entirely in the browser on your own hardware.


This post is part of GitHub Tools, where every repository is cloned and scanned before it is written up. The scan is a snapshot of one commit on one day; the repository has moved on since, so check it before you install.