Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
PythonPaddlePaddle/PaddleOCR

PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

84.7/100
89.4KForks: 11.3K
View on GitHubHomepage →
Loading report...

Similar Projects

ade-cli

80

The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal

Python2.4K

ParseBench

78

ParseBench - A Document Parsing Benchmark for AI Agents

Python576

MinerU

94

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

Python79.8K

paperless-ngx

93

A community-supported supercharged document management system: scan, index and archive all your documents

Python45.1K
Back to List