A llamafile is a language model and the program that runs it, packed into a single file. Download it, mark it executable, and run it: the same file works on macOS, Linux, Windows and the BSDs, on both x86 and Arm, with no installer, no Python and no separate model download. It starts a chat in the terminal and a local server with a web interface, built on llama.cpp. Whisperfile does the same for speech-to-text with whisper.cpp.
The trick is Cosmopolitan Libc, which builds one binary that each operating system recognizes as its own. That makes llamafile the easiest way to hand someone a model that will just run, or to archive one that will still run years from now. The 0.10 series moved to a new build system that tracks upstream llama.cpp more closely, so recent model families work; older 0.9 releases remain available for anyone who preferred the classic version.
llamafile began as a Mozilla Builders project, created by Justine Tunney, the author of Cosmopolitan, and is now maintained by Mozilla.ai. It has about 26,000 stars; the project is Apache-2.0, and its changes to llama.cpp and whisper.cpp are MIT, like those projects.
- Repository: github.com/mozilla-ai/llamafile
- Licence: custom (Other)
- Language: C++. Stars: 26.2K. Forks: 1,635. Last push: Oct 1, 2026.
- Scan: safe, Sep 30, 2026, commit fef6e44
Who it is for
Anyone who wants to try a local model with the fewest possible steps, people sharing a model with non-technical colleagues, and developers who need a portable offline model server.
Getting started
1. Download a small example model (Qwen3.5 0.8B)
curl -LO https://huggingface.co/mozilla-ai/llamafile_0.10/resolve/main/Qwen3.5-0.8B-Q8_0.llamafile2. Make it executable (macOS, Linux and BSD; on Windows, rename it to end in .exe instead)
chmod +x Qwen3.5-0.8B-Q8_0.llamafile3. Run it
./Qwen3.5-0.8B-Q8_0.llamafileWindows cannot run executables over 4 GB, so larger llamafiles will not start there; download the plain llamafile binary from the releases page and point it at a separate GGUF model instead. Larger, more capable prebuilt llamafiles are listed in the documentation.
Safety scan
We cloned mozilla-ai/llamafile at commit fef6e44 on Sep 30, 2026 and ran the checks described on the GitHub Tools page: credential patterns, decode-and-execute code, install-time scripts, committed binaries, risky CI workflows, every host the code talks to, known vulnerabilities in pinned dependencies, and project hygiene. A person read every hit. This is what we found.
- All 11 secret hits are in third_party/mbedtls (certs.c, pkparse.c, pkwrite.c): the test certificates and PEM header constants that ship with mbed TLS for its self-tests, public in that project. None is a credential. No pattern hits, install hooks or bare-IP URLs across 728 files and about 58,000 lines of C and C++.
- The one committed binary, build/gperf (750 KB), is a prebuilt copy of gperf, the GNU perfect-hash generator, used during the build. It is detected as a Windows executable because Cosmopolitan's portable binaries start with the same header.
- Hosts in the code are licence links, Mozilla's docs, Hugging Face, justine.lol and localscore.ai; we found no telemetry.
- There are no lockfiles: llama.cpp, whisper.cpp, mbed TLS and the other dependencies are vendored under third_party, so no advisory count, and their fixes arrive with llamafile releases.
- 6 workflows. labeler.yml uses pull_request_target but checks out the base branch and runs actions/labeler; the one third-party action is pinned to a commit. Security policy, licence and contributing guide present; no Dependabot or CodeQL.
What the scanner counted
| Check | Result |
|---|---|
| Secrets | 11 candidates found and read; see the notes above. |
| Suspicious code | None found. |
| Install-time code | None: nothing runs at install beyond the package manager itself. |
| Committed binaries | 1 executable or compiled object committed; listed under the raw findings. |
| CI workflows | 6 workflows. 1 uses pull_request_target, none check out the pull request head. 0 of 1 third-party action pinned to a tag rather than a commit. |
| Network hosts | 11 distinct hosts referenced from source; most often www.apache.org, docs.mozilla.ai, github.com, huggingface.co. No URLs to bare IP addresses. |
| Known vulnerabilities | No lockfile to check: dependencies are declared as ranges, so what gets installed is whatever is current on the day. |
| Project hygiene | Has security policy, licence file, contributing guide. Missing automated dependency updates, CodeQL. |
| OpenSSF Scorecard | Not scored: the project is not in Scorecard's weekly index. |
The raw findings
Every hit the scanner wrote out, with a link to the exact line at the scanned commit. Secrets candidates are redacted.
Secret candidates (11, redacted)
| Where | Rule | Match |
|---|---|---|
| third_party/mbedtls/certs.c:108 | private-key | -----B…--- (30 chars) |
| third_party/mbedtls/certs.c:347 | private-key | -----B…--- (31 chars) |
| third_party/mbedtls/certs.c:574 | private-key | -----B…--- (30 chars) |
| third_party/mbedtls/certs.c:801 | private-key | -----B…--- (31 chars) |
| third_party/mbedtls/certs.c:1017 | private-key | -----B…--- (30 chars) |
| third_party/mbedtls/certs.c:1146 | private-key | -----B…--- (31 chars) |
| third_party/mbedtls/pkparse.c:1248 | private-key | -----B…--- (31 chars) |
| third_party/mbedtls/pkparse.c:1279 | private-key | -----B…--- (30 chars) |
| third_party/mbedtls/pkparse.c:1309 | private-key | -----B…--- (27 chars) |
| third_party/mbedtls/pkwrite.c:497 | private-key | -----B…--- (31 chars) |
| third_party/mbedtls/pkwrite.c:499 | private-key | -----B…--- (30 chars) |
Committed binaries (1)
build/gperf: PE (Windows executable), 750 KB
Workflows worth a look
- .github/workflows/labeler.yml: pull_request_target
By the numbers
| Stars | 26.2K |
|---|---|
| Forks | 1,635 |
| Contributors | 77 |
| Commits | 869 |
| Open issues | 190 |
| Open pull requests | 24 |
| Releases | 43 |
| Latest release | 0.10.6 |
| Licence | custom |
| Main language | C++ |
| Project age | 3 years |
| Last push | Oct 1, 2026 |
| Tracked files | 728 |
| Lines of code | 58.4K |
| Checkout size | 33 MB |
Lines by language: C++ 26.6K, C 7,539, Markdown 5,645, C/C++ header 4,951, Python 3,571, Make 2,884.
Questions
Is llamafile free?
Yes. llamafile is Apache-2.0, with its changes to llama.cpp and whisper.cpp under MIT, and is free for any use. The models packed into llamafiles carry their own licences, which you should check before using one commercially.
What is the difference between llamafile and Ollama?
Both run models locally on llama.cpp. Ollama is a background service with a model registry and a command set you install once. llamafile needs no installation: the model and runtime are one file you can copy anywhere and run, which suits sharing, offline use and archiving.
Can I make my own llamafile?
Yes. The documentation's Creating llamafiles section shows how to combine the llamafile runtime with any GGUF model and default settings into one executable. You can also run the llamafile binary directly against a separate GGUF file without packing them together.
This post is part of GitHub Tools, where every repository is cloned and scanned before it is written up. The scan is a snapshot of one commit on one day; the repository has moved on since, so check it before you install.
