r/Rag
I benchmarked 7 document parsing APIs on the same 11 PDFs. None of them were good at everything.
- upvotes
- 24
- comments
- 19
Post
Highlighted: the lines this signal was extracted from
I got tired of document parsing comparisons that basically compare pricing pages, so I actually called the APIs. I took 11 documents from public datasets, including invoices, a photographed receipt, contracts, a French bank statement, a noisy scanned form, a handwritten cheque, and a medical EOB. 35 pages total. 119 API calls. Every provider got the same PDF, same JSON schema, same extraction instructions, same timeout, and the normal/default mode. I also deliberately included fields where the correct answer was null to see which systems would guess anyway. Here’s where things landed: Provider Field accuracy Row F1 Hallucinations Missing Median latency Claude 0.991 0.99 0 2 6.4s GPT 0.982 0.99 0 2 10.1s Reducto 0.982 0.99 0 4 10.1s Extend 0.962 1.00 0 2 22.1s Textract 0.936 0.99 2 1 14.4s LlamaExtract 0.903 0.99 3 3 22.5s Mistral OCR 0.884 0.99 3 3 4.2s There wasn’t really one winner here, which was probably the most useful part of the test. Claude had the best raw field accuracy. GPT was the cheapest per correct field. Mistral was the fastest. Extend was the only one that got 1.00 row F1, so it didn’t miss a single table row in this run. It also had zero hallucinations, though Claude, GPT, and Reducto did too. Here's some random stuff I noticed while going through the failures. Mistral and Textract would sometimes see...
Keep reading with a free account
The rest of this post, and every signal for Mistral AI, is in your free account.
Also quoted as evidence
Mistral and Textract would sometimes see an empty field and grab some other value from the document instead.
From the post
Mistral OCR Field accuracy 0.884 Row F1 0.99 Hallucinations 3 Missing 3 Median latency 4.2s
From the post