← Back to discoveryVERIFICATIONSnapshot first. Upstream always wins for current facts. REGISTRY PROVENANCE
REGISTRY SNAPSHOTSOURCE: GITHUB
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
View upstream on GitHub ↗50k+ star band
- Primary language
- Python
- Stars
- 89,920
- Forks
- 11,375
- Open issues
- 251
- License
- Apache-2.0
- Activity
- active
- Star band
- 50k+
- Archived
- No
ai4sciencechineseocrdocument-parsingdocument-translationkieocrpaddleocr-vlpdf-extractor-ragpdf-parserpdf2markdownpp-ocrpp-structurerag
This record reflects the RepoSource Registry snapshot generated at 2026-09-21T08:01:18Z. Mutable repository fields can change on GitHub after collection.
Source github-public-apiDataset 1.0.0Schema 1.1.0Indexed 2026-09-21T08:01:18Z