5 min read

PaddleOCR: Turn Images and PDFs Into Structured Text (GitHub, Scanned)

Baidu's OCR toolkit that reads text in 100+ languages and turns documents into Markdown or JSON.

PaddleOCR logo
✅
Scan: safe. Nothing to warn about in the core package; the known advisories sit in side projects (a LangChain integration, a browser SDK and a TypeScript API client), not in what pip install paddleocr installs. Scanned Sep 16, 2026; the full report is below.

PaddleOCR reads text out of images and PDFs. At its simplest, one command finds every line of text in a photo or scan and returns the words with their positions; the PP-OCRv6 models handle 50 languages in a single model, and the toolkit covers more than 100 in all, along with street signs, ID cards, dot-matrix print and other hard cases. Beyond plain OCR, PP-StructureV3 and the PaddleOCR-VL vision-language models parse whole documents, including tables, formulas, charts and seals, into Markdown or JSON ready to feed to a language model.

It is on this list because it has become plumbing for AI apps: Dify, RAGFlow and Cherry Studio use it to get documents into retrieval systems. The models are small, from 1.5 million parameters for the tiny OCR model to 0.9 billion for PaddleOCR-VL, so they run on a CPU or edge device, and they can be exported to ONNX or run through OpenVINO, TensorRT or Hugging Face Transformers instead of the PaddlePaddle framework.

PaddleOCR is developed by the PaddlePaddle team at Baidu, has about 90,000 stars and is licensed Apache-2.0. Version 3.7.0, with PP-OCRv6, was released in June 2026.

Who it is for

Developers building document pipelines, RAG systems or data-entry automation, and anyone who needs accurate OCR for languages beyond English, especially Chinese and Japanese.

Getting started

1. Install the PaddlePaddle framework, CPU build shown (the docs list CUDA builds)

python -m pip install paddlepaddle==3.2.0 -i https://www.paddlepaddle.org.cn/packages/stable/cpu/

2. Install PaddleOCR (use "paddleocr[all]" for document parsing and translation too)

python -m pip install paddleocr

3. Read the text in a sample image and save the results

paddleocr ocr -i https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/general_ocr_002.png --save_path ./output

4. Parse a whole document page into Markdown and JSON with PP-StructureV3

paddleocr pp_structurev3 -i ./your_document.png

Models download on first use. To try it without installing anything, paddleocr.com has an online Experience Center and APIs.

Safety scan

We cloned PaddlePaddle/PaddleOCR at commit dab3fe3 on Sep 16, 2026 and ran the checks described on the GitHub Tools page: credential patterns, decode-and-execute code, install-time scripts, committed binaries, risky CI workflows, every host the code talks to, known vulnerabilities in pinned dependencies, and project hygiene. A person read every hit. This is what we found.

  • No secrets, no pattern hits, no bare-IP URLs and no install hooks across 2,455 files and about 359,000 lines (121,000 Python, 150,000 Markdown). The two binaries are Gradle wrapper jars in the Android demo apps. The installer scripts are demos; the Arm Virtual Hardware one runs sudo pip install and a sudo setup script on that virtual board.
  • Network hosts are Baidu's object storage (paddleocr.bj.bcebos.com and paddle-model-ecology.bj.bcebos.com) for model weights and demo images, GitHub, and Apache licence links in file headers; we found no analytics.
  • The paddleocr package has no lockfile, so its own dependencies were not checked. The 131 advisories (3 critical, 60 high) are in langchain-paddleocr/uv.lock (73, including anyio, langchain-core, langsmith and pillow), paddleocr-js (49) and api_sdk/typescript (9, build tooling). In the browser SDK, protobufjs 7.5.4 (critical) is a runtime dependency through onnxruntime-web, which uses it to read model files; vitest's critical is test tooling.
  • 9 workflows, none using pull_request_target; all 4 third-party actions are pinned to commits. Dependabot and licence present; no security policy, contributing guide or CodeQL. OpenSSF Scorecard 4.2.

What the scanner counted

CheckResult
SecretsNone found.
Suspicious codeNone found.
Install-time code3 installer scripts (one can call sudo)
Committed binaries2 executable or compiled objects committed; listed under the raw findings.
CI workflows9 workflows. None use pull_request_target. 0 of 4 third-party actions pinned to a tag rather than a commit.
Network hosts40 distinct hosts referenced from source; most often www.apache.org, paddleocr.bj.bcebos.com, github.com, paddle-model-ecology.bj.bcebos.com. No URLs to bare IP addresses.
Known vulnerabilities131 advisories across 634 pinned packages: 3 critical, 60 high, 53 moderate, 15 low. api_sdk/typescript/package-lock.json: 166 packages, 9 advisories; deploy/paddleocr_vl_docker/hps/gateway/requirements.txt: 2 packages, 0 advisories; docs/version2.x/algorithm/formula_recognition/requirements.txt: 1 packages, 0 advisories; langchain-paddleocr/uv.lock: 112 packages, 73 advisories; paddleocr-js/package-lock.json: 390 packages, 49 advisories; ppstructure/kie/requirements.txt: 1 packages, 0 advisories; test_tipc/supplementary/requirements.txt: 1 packages, 0 advisories.
Project hygieneHas automated dependency updates, licence file. Missing security policy, CodeQL, contributing guide.
OpenSSF Scorecard4.2 out of 10, as of Sep 28, 2026.

The raw findings

Every hit the scanner wrote out, with a link to the exact line at the scanned commit. Secrets candidates are redacted.

Installer scripts (3)
Committed binaries (2)
  • deploy/android_demo/gradle/wrapper/gradle-wrapper.jar: JAR, 54 KB
  • deploy/ppocr-android/gradle/wrapper/gradle-wrapper.jar: JAR, 44 KB
Worst known vulnerabilities (24 of 131)
AdvisorySeverityPackageSummary
GHSA-82r6-8w77-94w6criticalanyio@4.12.1AnyIO: TLSStream IDNA 2003 host name encoding enables potential TLS certificate spoofing
GHSA-xq3m-2v4x-88ggcriticalprotobufjs@7.5.4Arbitrary code execution in protobufjs
GHSA-5xrq-8626-4rwpcriticalvitest@3.2.4When Vitest UI server is listening, arbitrary file can be read and executed
GHSA-28wg-ghj8-5hjvhighnanoid@3.3.12nanoid: non-secure generators can loop indefinitely with negative size
GHSA-2v37-7h3g-55p8highnanoid@3.3.12nanoid: custom generators can loop indefinitely when size is zero
GHSA-r28c-9q8g-f849highpostcss@8.5.15PostCSS: Path Traversal in Previous Source Map Auto-Loading (sourceMappingURL) leads to Arbitrary .map File Disclosure
GHSA-fx2h-pf6j-xcffhighvite@8.0.14vite: `server.fs.deny` bypass on Windows alternate paths
GHSA-cq5v-8q36-5273highaiohttp@3.13.3AIOHTTP: Out-of-bounds heap read in C HTTP response parser error path (malformed chunked response)
GHSA-47fr-3ffg-hgmwhighclick@8.3.1
GHSA-pjwx-r37v-7724highlangchain-core@1.2.7LangChain vulnerable to unsafe deserialization of attacker-controlled objects through overly broad `load()` allowlists
GHSA-qh6h-p6c9-ff54highlangchain-core@1.2.7LangChain Core has Path Traversal vulnerabilites in legacy `load_prompt` functions
GHSA-3644-q5cj-c5c7highlangsmith@0.6.7LangSmith SDK: Public prompt pull deserializes untrusted manifests without trust boundary warning
GHSA-f4xh-w4cj-qxq8highlangsmith@0.6.7LangSmith SDK TracingMiddleware: Arbitrary server-side file read
GHSA-45hq-cxwh-f6vchighpillow@12.1.0Pillow `BdfFontFile`: `Image.new()` called without `_decompression_bomb_check()` - bomb protection bypass via font loadi…
GHSA-5x94-69rx-g8h2highpillow@12.1.0Pillow: `FontFile.compile()`: `Image.new()` called without `_decompression_bomb_check()`
GHSA-62p4-gmf7-7g93highpillow@12.1.0Pillow: Out-of-bounds read via attacker-controlled row stride on Pillow's mmap path (McIdas AREA files)
GHSA-6r8x-57c9-28j4highpillow@12.1.0Pillow: Heap out-of-bounds write `Image.paste()` / `Image.crop()` via signed coordinate overflow
GHSA-8v84-f9pq-wr9xhighpillow@12.1.0Pillow `PcfFontFile._load_bitmaps()`: `Image.frombytes()` called without `_decompression_bomb_check()` - bomb protection…
GHSA-9hw9-ch79-4vh6highpillow@12.1.0Pillow: Controlled heap out-of-bounds write in Pillow `ImageCmsTransform.apply()` via output mode mismatch
GHSA-cfh3-3jmp-rvhchighpillow@12.1.0Pillow affected by out-of-bounds write when loading PSD images
GHSA-jjj6-mw9f-p565highpillow@12.1.0Pillow: Decompression Bomb DoS via PdfParser.PdfStream.decode()
GHSA-phj9-mv4w-65pmhighpillow@12.1.0Pillow `GdImageFile._open()`: image dimensions accepted without `_decompression_bomb_check()`
GHSA-pwv6-vv43-88grhighpillow@12.1.0Pillow has an OOB Write with Invalid PSD Tile Extents (Integer Overflow)
GHSA-vjc4-5qp5-m44jhighpillow@12.1.0Pillow JPEG2000 tiled decode retains a growing scratch buffer and can be used for denial of service

By the numbers

Stars90.5K
Forks11.4K
Contributors295
Commits6,927
Open issues159
Open pull requests77
Releases33
Latest releasev3.7.0
LicenceApache-2.0
Main languagePython
Project age6 years
Last pushSep 16, 2026
Tracked files2,455
Lines of code358.9K
Checkout size136 MB

Lines by language: Markdown 149.6K, Python 121K, YAML 28.5K, C++ 20.8K, TypeScript 11.9K, Shell 6,774.

Questions

Is PaddleOCR free?

Yes. PaddleOCR and its models are Apache-2.0, free for commercial use. Baidu also offers an online Experience Center and API at paddleocr.com for people who would rather not run it themselves.

Does PaddleOCR need a GPU?

No. The PP-OCR models are small enough for a CPU; the 3.7.0 release notes cite a 5.2 times CPU speedup with OpenVINO. A GPU helps for large batches and for the PaddleOCR-VL document model.

Can I use PaddleOCR without the PaddlePaddle framework?

Yes. Since version 3.5 you can choose the inference engine: PaddlePaddle, Hugging Face Transformers (for about 20 major models) or ONNX Runtime with exported models. The docs recommend installing only one engine per environment, since their dependencies can conflict.


This post is part of GitHub Tools, where every repository is cloned and scanned before it is written up. The scan is a snapshot of one commit on one day; the repository has moved on since, so check it before you install.