PaddleOCR reads text out of images and PDFs. At its simplest, one command finds every line of text in a photo or scan and returns the words with their positions; the PP-OCRv6 models handle 50 languages in a single model, and the toolkit covers more than 100 in all, along with street signs, ID cards, dot-matrix print and other hard cases. Beyond plain OCR, PP-StructureV3 and the PaddleOCR-VL vision-language models parse whole documents, including tables, formulas, charts and seals, into Markdown or JSON ready to feed to a language model.
It is on this list because it has become plumbing for AI apps: Dify, RAGFlow and Cherry Studio use it to get documents into retrieval systems. The models are small, from 1.5 million parameters for the tiny OCR model to 0.9 billion for PaddleOCR-VL, so they run on a CPU or edge device, and they can be exported to ONNX or run through OpenVINO, TensorRT or Hugging Face Transformers instead of the PaddlePaddle framework.
PaddleOCR is developed by the PaddlePaddle team at Baidu, has about 90,000 stars and is licensed Apache-2.0. Version 3.7.0, with PP-OCRv6, was released in June 2026.
- Repository: github.com/PaddlePaddle/PaddleOCR
- Licence: Apache-2.0 (Apache License 2.0)
- Language: Python. Stars: 90.5K. Forks: 11.4K. Last push: Sep 16, 2026.
- Scan: safe, Sep 16, 2026, commit dab3fe3
Who it is for
Developers building document pipelines, RAG systems or data-entry automation, and anyone who needs accurate OCR for languages beyond English, especially Chinese and Japanese.
Getting started
1. Install the PaddlePaddle framework, CPU build shown (the docs list CUDA builds)
python -m pip install paddlepaddle==3.2.0 -i https://www.paddlepaddle.org.cn/packages/stable/cpu/2. Install PaddleOCR (use "paddleocr[all]" for document parsing and translation too)
python -m pip install paddleocr3. Read the text in a sample image and save the results
paddleocr ocr -i https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/general_ocr_002.png --save_path ./output4. Parse a whole document page into Markdown and JSON with PP-StructureV3
paddleocr pp_structurev3 -i ./your_document.pngModels download on first use. To try it without installing anything, paddleocr.com has an online Experience Center and APIs.
Safety scan
We cloned PaddlePaddle/PaddleOCR at commit dab3fe3 on Sep 16, 2026 and ran the checks described on the GitHub Tools page: credential patterns, decode-and-execute code, install-time scripts, committed binaries, risky CI workflows, every host the code talks to, known vulnerabilities in pinned dependencies, and project hygiene. A person read every hit. This is what we found.
- No secrets, no pattern hits, no bare-IP URLs and no install hooks across 2,455 files and about 359,000 lines (121,000 Python, 150,000 Markdown). The two binaries are Gradle wrapper jars in the Android demo apps. The installer scripts are demos; the Arm Virtual Hardware one runs sudo pip install and a sudo setup script on that virtual board.
- Network hosts are Baidu's object storage (paddleocr.bj.bcebos.com and paddle-model-ecology.bj.bcebos.com) for model weights and demo images, GitHub, and Apache licence links in file headers; we found no analytics.
- The paddleocr package has no lockfile, so its own dependencies were not checked. The 131 advisories (3 critical, 60 high) are in langchain-paddleocr/uv.lock (73, including anyio, langchain-core, langsmith and pillow), paddleocr-js (49) and api_sdk/typescript (9, build tooling). In the browser SDK, protobufjs 7.5.4 (critical) is a runtime dependency through onnxruntime-web, which uses it to read model files; vitest's critical is test tooling.
- 9 workflows, none using pull_request_target; all 4 third-party actions are pinned to commits. Dependabot and licence present; no security policy, contributing guide or CodeQL. OpenSSF Scorecard 4.2.
What the scanner counted
| Check | Result |
|---|---|
| Secrets | None found. |
| Suspicious code | None found. |
| Install-time code | 3 installer scripts (one can call sudo) |
| Committed binaries | 2 executable or compiled objects committed; listed under the raw findings. |
| CI workflows | 9 workflows. None use pull_request_target. 0 of 4 third-party actions pinned to a tag rather than a commit. |
| Network hosts | 40 distinct hosts referenced from source; most often www.apache.org, paddleocr.bj.bcebos.com, github.com, paddle-model-ecology.bj.bcebos.com. No URLs to bare IP addresses. |
| Known vulnerabilities | 131 advisories across 634 pinned packages: 3 critical, 60 high, 53 moderate, 15 low. api_sdk/typescript/package-lock.json: 166 packages, 9 advisories; deploy/paddleocr_vl_docker/hps/gateway/requirements.txt: 2 packages, 0 advisories; docs/version2.x/algorithm/formula_recognition/requirements.txt: 1 packages, 0 advisories; langchain-paddleocr/uv.lock: 112 packages, 73 advisories; paddleocr-js/package-lock.json: 390 packages, 49 advisories; ppstructure/kie/requirements.txt: 1 packages, 0 advisories; test_tipc/supplementary/requirements.txt: 1 packages, 0 advisories. |
| Project hygiene | Has automated dependency updates, licence file. Missing security policy, CodeQL, contributing guide. |
| OpenSSF Scorecard | 4.2 out of 10, as of Sep 28, 2026. |
The raw findings
Every hit the scanner wrote out, with a link to the exact line at the scanned commit. Secrets candidates are redacted.
Installer scripts (3)
- deploy/avh/run_demo.sh, 185 lines, uses sudo; talks to paddleocr.bj.bcebos.com, www.apache.org
- deploy/ios_demo/scripts/run_benchmark.sh, 897 lines; talks to www.apache.org
- deploy/ppocr-android/run_benchmark.sh, 40 lines
Committed binaries (2)
deploy/android_demo/gradle/wrapper/gradle-wrapper.jar: JAR, 54 KBdeploy/ppocr-android/gradle/wrapper/gradle-wrapper.jar: JAR, 44 KB
Worst known vulnerabilities (24 of 131)
| Advisory | Severity | Package | Summary |
|---|---|---|---|
| GHSA-82r6-8w77-94w6 | critical | anyio@4.12.1 | AnyIO: TLSStream IDNA 2003 host name encoding enables potential TLS certificate spoofing |
| GHSA-xq3m-2v4x-88gg | critical | protobufjs@7.5.4 | Arbitrary code execution in protobufjs |
| GHSA-5xrq-8626-4rwp | critical | vitest@3.2.4 | When Vitest UI server is listening, arbitrary file can be read and executed |
| GHSA-28wg-ghj8-5hjv | high | nanoid@3.3.12 | nanoid: non-secure generators can loop indefinitely with negative size |
| GHSA-2v37-7h3g-55p8 | high | nanoid@3.3.12 | nanoid: custom generators can loop indefinitely when size is zero |
| GHSA-r28c-9q8g-f849 | high | postcss@8.5.15 | PostCSS: Path Traversal in Previous Source Map Auto-Loading (sourceMappingURL) leads to Arbitrary .map File Disclosure |
| GHSA-fx2h-pf6j-xcff | high | vite@8.0.14 | vite: `server.fs.deny` bypass on Windows alternate paths |
| GHSA-cq5v-8q36-5273 | high | aiohttp@3.13.3 | AIOHTTP: Out-of-bounds heap read in C HTTP response parser error path (malformed chunked response) |
| GHSA-47fr-3ffg-hgmw | high | click@8.3.1 | |
| GHSA-pjwx-r37v-7724 | high | langchain-core@1.2.7 | LangChain vulnerable to unsafe deserialization of attacker-controlled objects through overly broad `load()` allowlists |
| GHSA-qh6h-p6c9-ff54 | high | langchain-core@1.2.7 | LangChain Core has Path Traversal vulnerabilites in legacy `load_prompt` functions |
| GHSA-3644-q5cj-c5c7 | high | langsmith@0.6.7 | LangSmith SDK: Public prompt pull deserializes untrusted manifests without trust boundary warning |
| GHSA-f4xh-w4cj-qxq8 | high | langsmith@0.6.7 | LangSmith SDK TracingMiddleware: Arbitrary server-side file read |
| GHSA-45hq-cxwh-f6vc | high | pillow@12.1.0 | Pillow `BdfFontFile`: `Image.new()` called without `_decompression_bomb_check()` - bomb protection bypass via font loadi… |
| GHSA-5x94-69rx-g8h2 | high | pillow@12.1.0 | Pillow: `FontFile.compile()`: `Image.new()` called without `_decompression_bomb_check()` |
| GHSA-62p4-gmf7-7g93 | high | pillow@12.1.0 | Pillow: Out-of-bounds read via attacker-controlled row stride on Pillow's mmap path (McIdas AREA files) |
| GHSA-6r8x-57c9-28j4 | high | pillow@12.1.0 | Pillow: Heap out-of-bounds write `Image.paste()` / `Image.crop()` via signed coordinate overflow |
| GHSA-8v84-f9pq-wr9x | high | pillow@12.1.0 | Pillow `PcfFontFile._load_bitmaps()`: `Image.frombytes()` called without `_decompression_bomb_check()` - bomb protection… |
| GHSA-9hw9-ch79-4vh6 | high | pillow@12.1.0 | Pillow: Controlled heap out-of-bounds write in Pillow `ImageCmsTransform.apply()` via output mode mismatch |
| GHSA-cfh3-3jmp-rvhc | high | pillow@12.1.0 | Pillow affected by out-of-bounds write when loading PSD images |
| GHSA-jjj6-mw9f-p565 | high | pillow@12.1.0 | Pillow: Decompression Bomb DoS via PdfParser.PdfStream.decode() |
| GHSA-phj9-mv4w-65pm | high | pillow@12.1.0 | Pillow `GdImageFile._open()`: image dimensions accepted without `_decompression_bomb_check()` |
| GHSA-pwv6-vv43-88gr | high | pillow@12.1.0 | Pillow has an OOB Write with Invalid PSD Tile Extents (Integer Overflow) |
| GHSA-vjc4-5qp5-m44j | high | pillow@12.1.0 | Pillow JPEG2000 tiled decode retains a growing scratch buffer and can be used for denial of service |
By the numbers
| Stars | 90.5K |
|---|---|
| Forks | 11.4K |
| Contributors | 295 |
| Commits | 6,927 |
| Open issues | 159 |
| Open pull requests | 77 |
| Releases | 33 |
| Latest release | v3.7.0 |
| Licence | Apache-2.0 |
| Main language | Python |
| Project age | 6 years |
| Last push | Sep 16, 2026 |
| Tracked files | 2,455 |
| Lines of code | 358.9K |
| Checkout size | 136 MB |
Lines by language: Markdown 149.6K, Python 121K, YAML 28.5K, C++ 20.8K, TypeScript 11.9K, Shell 6,774.
Questions
Is PaddleOCR free?
Yes. PaddleOCR and its models are Apache-2.0, free for commercial use. Baidu also offers an online Experience Center and API at paddleocr.com for people who would rather not run it themselves.
Does PaddleOCR need a GPU?
No. The PP-OCR models are small enough for a CPU; the 3.7.0 release notes cite a 5.2 times CPU speedup with OpenVINO. A GPU helps for large batches and for the PaddleOCR-VL document model.
Can I use PaddleOCR without the PaddlePaddle framework?
Yes. Since version 3.5 you can choose the inference engine: PaddlePaddle, Hugging Face Transformers (for about 20 major models) or ONNX Runtime with exported models. The docs recommend installing only one engine per environment, since their dependencies can conflict.
This post is part of GitHub Tools, where every repository is cloned and scanned before it is written up. The scan is a snapshot of one commit on one day; the repository has moved on since, so check it before you install.
