PDFMathTranslate, run from the command line as pdf2zh, translates a PDF while leaving its layout alone. Give it a paper and it writes two files: a translated version and a bilingual one with the original and translation page by page, with formulas, charts, tables of contents and annotations kept where they were. A layout model (DocLayout-YOLO, run through ONNX) works out which parts of each page are text and which are equations or figures, so only the prose is sent for translation.
It is on this list because translating academic papers is the job most PDF translators do badly, and this one was built for it. It works with Google Translate by default and with DeepL, OpenAI, Ollama and many other services, and comes as a Python package, a browser GUI, a Docker image, a Windows build and a Zotero plugin. Recent additions include an ultrafast mode for PDFs that already contain text and experimental OCR for scanned pages.
The project has about 37,000 stars, is licensed AGPL-3.0, and was presented as a demo paper at EMNLP 2025. Version 2.0, built on the BabelDOC engine, lives in a separate repository, PDFMathTranslate-next; this repository is the 1.x line, still updated.
- Repository: github.com/PDFMathTranslate/PDFMathTranslate
- Licence: AGPL-3.0 (GNU Affero General Public License v3.0)
- Language: Python. Stars: 37.3K. Forks: 3,339. Last push: Oct 3, 2026.
- Scan: safe, Oct 3, 2026, commit fd91ca0
Who it is for
Students and researchers who read papers in a language they are less fluent in, and anyone who has watched a regular translator destroy the equations and columns of a technical PDF.
Getting started
1. Install with uv (Python 3.11 or 3.12; plain pip install pdf2zh also works)
pip install uv && uv tool install --python 3.12 pdf2zh2. Translate a PDF into English and Chinese files in the current folder (Google is the default service)
pdf2zh document.pdf3. Pick languages and a service, here English to Japanese through DeepL
pdf2zh document.pdf -li en -lo ja -s deepl4. Or run the browser interface, which opens on localhost:7860
pdf2zh -iThe first run downloads the layout model from Hugging Face. Services other than Google and Bing need your own API key, and anything you translate is sent to the service you pick unless you use a local model through Ollama.
Safety scan
We cloned PDFMathTranslate/PDFMathTranslate at commit fd91ca0 on Oct 3, 2026 and ran the checks described on the GitHub Tools page: credential patterns, decode-and-execute code, install-time scripts, committed binaries, risky CI workflows, every host the code talks to, known vulnerabilities in pinned dependencies, and project hygiene. A person read every hit. This is what we found.
- No secrets, no pattern hits, no bare-IP URLs, no install-time hooks and no committed binaries across 88 files and about 13,000 lines, 8,300 of them Python.
- Every external host in the code is a translation backend you can select (Google, Bing, DeepL, DeepLX, OpenAI, Azure, Gemini, x.ai, Zhipu, SiliconFlow, 302.AI, ModelScope and others) or a project link. The text of your PDF goes to whichever one you pick, and stays on your machine only with a local model through Ollama.
- The layout model it downloads from Hugging Face (DocLayout-YOLO) is in ONNX format, which is a graph of operations rather than a pickled Python object, so loading it cannot run arbitrary code.
- There is no lockfile: pyproject.toml declares version ranges, so there were no pinned dependencies for OSV to check, and what you get depends on what pip resolves on the day. Installing with uv tool install keeps it in its own environment.
- 7 workflows, none using pull_request_target; 20 of 29 third-party actions are pinned to a commit. Dependabot and licence present; no security policy, contributing guide or CodeQL.
What the scanner counted
| Check | Result |
|---|---|
| Secrets | None found. |
| Suspicious code | None found. |
| Install-time code | None: nothing runs at install beyond the package manager itself. |
| Committed binaries | None. |
| CI workflows | 7 workflows. None use pull_request_target. 9 of 29 third-party actions pinned to a tag rather than a commit. |
| Network hosts | 34 distinct hosts referenced from source; most often github.com, raw.githubusercontent.com, api.openailiked.com, aka.ms. No URLs to bare IP addresses. |
| Known vulnerabilities | No lockfile to check: dependencies are declared as ranges, so what gets installed is whatever is current on the day. |
| Project hygiene | Has automated dependency updates, licence file. Missing security policy, CodeQL, contributing guide. |
| OpenSSF Scorecard | Not scored: the project is not in Scorecard's weekly index. |
By the numbers
| Stars | 37.3K |
|---|---|
| Forks | 3,339 |
| Contributors | 59 |
| Commits | 2,369 |
| Open issues | 130 |
| Open pull requests | 38 |
| Releases | 24 |
| Latest release | v1.9.11 |
| Licence | AGPL-3.0 |
| Main language | Python |
| Project age | 2 years |
| Last push | Oct 3, 2026 |
| Tracked files | 88 |
| Lines of code | 13K |
| Checkout size | 7 MB |
Lines by language: Python 8,337, Markdown 2,791, YAML 1,633, TOML 104, PowerShell 81, Batch 28.
Questions
Is PDFMathTranslate free?
Yes. The software is AGPL-3.0 and free, and its default service, Google Translate, needs no key. Paid services such as DeepL or OpenAI bill you through your own API key. The project also runs a free online version at pdf2zh.com with limited capacity, and Immersive Translate offers a hosted BabelDOC version with a free quota.
Does PDFMathTranslate work on scanned PDFs?
Partly. Experimental OCR, installed with pip install 'pdf2zh[ocr]', reads image-only pages before translating them. It targets clean white-background scans; handwriting and inline equations may be misread, and pages that already mix text and scans are skipped.
What is the difference between PDFMathTranslate and PDFMathTranslate-next?
PDFMathTranslate-next is version 2.0, moved to its own repository and built on the BabelDOC engine. This repository is the 1.x line, which remains maintained and can call the 2.0 engine experimentally with --mode precise or --babeldoc.
This post is part of GitHub Tools, where every repository is cloned and scanned before it is written up. The scan is a snapshot of one commit on one day; the repository has moved on since, so check it before you install.
