r/Rag
Best open source OCR models to replace Textract (tested a bunch)
- upvotes
- 32
- comments
- 22
Post
Been testing self-hosted OCR to get off Textract, since the managed services sit around $1.50 per 1,000 pages and jump hard once forms and tables are involved. The main lesson is there's no single winner, it comes down to your document type. What I tried: docling: turns digital PDFs into clean markdown, great for RAG ingestion, shaky on scans and handwriting PaddleOCR-VL: multilingual VLM, 94.5% on OmniDocBench v1.5, small at ~0.9B so it fits a modest GPU MinerU: best of the bunch on formula-heavy scientific PDFs GLM-OCR: solid general image-to-markdown, good baseline to measure the rest against Two things that caught me out: the top scores are close enough that testing on your own documents matters more than the leaderboard, and licenses bite, since Surya needs a commercial license past a revenue threshold and a couple of the leaders are CC-BY-NC. If you'd rather not run a separate server per model, some inference servers like SIE let you swap OCR models behind one endpoint, but that's optional.
Extracted from these lines
[comment u/jc-atg] From our experience, Mistral OCR (not open source) is the best solution, but expensive ($2-$4/1k pages). PaddleOCR is almost as good, LightOnOCR was far behind in our testings (in most cases, Docling was better).