← Back to discovery
REGISTRY SNAPSHOTSOURCE: GITHUB

PaddlePaddle/PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

Primary language
Python
Stars
89,920
Forks
11,375
Open issues
251
License
Apache-2.0
Activity
active
Star band
50k+
Archived
No
ai4sciencechineseocrdocument-parsingdocument-translationkieocrpaddleocr-vlpdf-extractor-ragpdf-parserpdf2markdownpp-ocrpp-structurerag
VERIFICATIONSnapshot first. Upstream always wins for current facts.

This record reflects the RepoSource Registry snapshot generated at 2026-09-21T08:01:18Z. Mutable repository fields can change on GitHub after collection.

REGISTRY PROVENANCE
Source github-public-apiDataset 1.0.0Schema 1.1.0Indexed 2026-09-21T08:01:18Z