Skip to main content
Mistral AICustomer feedback

A user from AskThisGuy.com reported that Mistral OCR is the best-performing solution they tested but is expensive at $2-$4 per 1,000 pages and is not open source.

What happened

Post: "Best open source OCR models to replace Textract (tested a bunch)"

Source

RedditSep 22, 2026By u/Green-Topic-1024

r/Rag

Best open source OCR models to replace Textract (tested a bunch)

upvotes
32
comments
22

Post

Been testing self-hosted OCR to get off Textract, since the managed services sit around $1.50 per 1,000 pages and jump hard once forms and tables are involved. The main lesson is there's no single winner, it comes down to your document type. What I tried: docling: turns digital PDFs into clean markdown, great for RAG ingestion, shaky on scans and handwriting PaddleOCR-VL: multilingual VLM, 94.5% on OmniDocBench v1.5, small at ~0.9B so it fits a modest GPU MinerU: best of the bunch on formula-heavy scientific PDFs GLM-OCR: solid general image-to-markdown, good baseline to measure the rest against Two things that caught me out: the top scores are close enough that testing on your own documents matters more than the leaderboard, and licenses bite, since Surya needs a commercial license past a revenue threshold and a couple of the leaders are CC-BY-NC. If you'd rather not run a separate server per model, some inference servers like SIE let you swap OCR models behind one endpoint, but that's optional.

Extracted from these lines

  • [comment u/jc-atg] From our experience, Mistral OCR (not open source) is the best solution, but expensive ($2-$4/1k pages). PaddleOCR is almost as good, LightOnOCR was far behind in our testings (in most cases, Docling was better).

reddit.com/r/Rag/comments/1wmyoiz/best_open_source_ocr_models_to_repl...Read the full source

Comments on the post

5 of 22 comments
  • “in my usecase mini-cpm-4 outperformed textract, only negative was, trust remote code = True in it's config. Customer didn't like that.”

    u/polandtown2 points · Sep 22, 2026View

  • “pp-ocrv6 is hands down the best replacement to AWS Textract if you're looking for true bounding boxes, and not markdown. PaddleOCR-VL, dots.mocr, GLM-OCR have a lot of variability in their outputs based on the domain, so I'd always recommend trying them out on your own data first. If you're interested in quick way to test (with free usage, in the anonymous tier), try: uvx vlmrun gw chat -m pa”

    u/vlm-run2 points · Sep 23, 2026View

  • “Side note since people are comparing models: the annoying part was never picking one of the options, it was running them side by side. I had four vLLM servers up at one point and the GPU was idle maybe 80% of the time while I stared at markdown diffs. Moved them behind SIE (Superlinked Inference Engine) for the last round (open source, full disclosure I follow the project). One endpoint, swap th”

    u/Green-Topic-10241 points · Sep 24, 2026View

  • “What are you scanning, just PDFs of reports etc? I was working for a long time on had filled worksheets and the LLM based ones were the best”

    u/GnarlyBear1 points · Sep 22, 2026View

  • “Have a look at LightonOCR, results on PDF, scans and handwriting was pretty good for me: https://lighton.ai/lighton-blogs/making-knowledge-machine-readable”

    u/Thick-Boat48961 points · Sep 22, 2026View

Extracted by Autobound

From the Signal API record
Signal
Customer feedback

What this signalsUser posts often show product pain before it reaches reviews or churn.

Subreddit
r/Rag

Companies

  • PaddleOCRAlso named
  • LightonAlso named

The full record

From the Signal API record

Numbers

Mentions
3

Details

Timing
Ongoing state
Category
Pricing
Virality
Somewhat high
Post kind
Text
Prominence
Aside
Company's role
Vendor

Topics and mentions

Topics

  • vendor evaluation
  • performance
  • ocr
  • pricing

Flair

  • Tools & Resources

Products named

  • Mistral OCR

Extraction

Sentiment
Mixed
Detected
Sep 22, 2026
signal_type
reddit-company
signal_subtype
customerFeedback

Use this data

Get every Reddit signal for Mistral AI and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at Mistral AI this week?”

  2. Send it to your own tools

    The Signal API returns Reddit signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full reddit-company record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/ef688fd4-ffba-52a0-aefd-40240948c3ba returns this record as JSON. POST /v1/companies/enrich returns every signal for mistral.ai.

{
  "signal_id": "ef688fd4-ffba-52a0-aefd-40240948c3ba",
  "signal_type": "reddit-company",
  "signal_subtype": "customerFeedback",
  "detected_at": "2026-09-22T03:45:45+00:00",
  "company": {
    "name": "Mistral AI",
    "domain": "mistral.ai"
  },
  "data": {
    "nsfw": false,
    "stage": "none",
    "awards": 0,
    "timing": "ongoing_state",
    "topics": [
      "ocr",
      "pricing",
      "vendor evaluation",
      "performance"
    ],
    "post_id": "1wmyoiz",
    "summary": "A user from AskThisGuy.com reported that Mistral OCR is the best-performing solution they tested but is expensive at $2-$4 per 1,000 pages and is not open source.",
    "category": "pricing",
    "comments": [
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wmyoiz/comment/pbdmt3x/",
        "depth": 0,
        "score": 2,
        "author": "polandtown",
        "excerpt": "in my usecase mini-cpm-4 outperformed textract, only negative was, trust remote code = True in it's config. Customer didn't like that.",
        "posted_at": "2026-09-22T15:04:21.000Z",
        "author_url": "https://www.reddit.com/user/polandtown/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wmyoiz/comment/pbnf0o1/",
        "depth": 0,
        "score": 2,
        "author": "vlm-run",
        "excerpt": "pp-ocrv6 is hands down the best replacement to AWS Textract if you're looking for true bounding boxes, and not markdown.\n\n PaddleOCR-VL, dots.mocr, GLM-OCR have a lot of variability in their outputs based on the domain, so I'd always recommend trying them out on your own data first.\n\n If you're interested in quick way to test (with free usage, in the anonymous tier), try:\n\nuvx vlmrun gw chat -m pa",
        "posted_at": "2026-09-23T21:36:58.000Z",
        "author_url": "https://www.reddit.com/user/vlm-run/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wmyoiz/comment/pbp53jz/",
        "depth": 0,
        "score": 1,
        "author": "Green-Topic-1024",
        "excerpt": "Side note since people are comparing models: the annoying part was never picking one of the options, it was running them side by side. I had four vLLM servers up at one point and the GPU was idle maybe 80% of the time while I stared at markdown diffs.\n\n Moved them behind SIE (Superlinked Inference Engine) for the last round (open source, full disclosure I follow the project). One endpoint, swap th",
        "posted_at": "2026-09-24T03:18:06.000Z",
        "author_url": "https://www.reddit.com/user/Green-Topic-1024/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wmyoiz/comment/pbbl2ik/",
        "depth": 0,
        "score": 1,
        "author": "GnarlyBear",
        "excerpt": "What are you scanning, just PDFs of reports etc? I was working for a long time on had filled worksheets and the LLM based ones were the best",
        "posted_at": "2026-09-22T07:16:34.000Z",
        "author_url": "https://www.reddit.com/user/GnarlyBear/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wmyoiz/comment/pbeap67/",
        "depth": 0,
        "score": 1,
        "author": "Thick-Boat4896",
        "excerpt": "Have a look at LightonOCR, results on PDF, scans and handwriting was pretty good for me: https://lighton.ai/lighton-blogs/making-knowledge-machine-readable",
        "posted_at": "2026-09-22T16:46:00.000Z",
        "author_url": "https://www.reddit.com/user/Thick-Boat4896/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wmyoiz/comment/pbj3lkn/",
        "depth": 0,
        "score": 1,
        "author": "jc-atg",
        "excerpt": "Hi, at www.askthisguy.com we tested Mistral OCR, PaddleOCR and LightOnOCR for many documents ; compared it with parsers like Docling, Markitdown, etc.\n\n From our experience, Mistral OCR (not open source) is the best solution, but expensive ($2-$4/1k pages). PaddleOCR is almost as good, LightOnOCR was far behind in our testings (in most cases, Docling was better).\n\n Generally speaking, do not trust",
        "posted_at": "2026-09-23T08:56:56.000Z",
        "author_url": "https://www.reddit.com/user/jc-atg/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wmyoiz/comment/pbc30ue/",
        "depth": 0,
        "score": 1,
        "author": "Anu_Rag9704",
        "excerpt": "There’s a new one by alibaba i guess?",
        "posted_at": "2026-09-22T09:53:32.000Z",
        "author_url": "https://www.reddit.com/user/Anu_Rag9704/"
      },
      {
        "url": "https://www.reddit.com/r/Rag/comments/1wmyoiz/comment/pbc8not/",
        "depth": 0,
        "score": 1,
        "author": "ephocalate",
        "excerpt": "but mineru isnt an ocr right? If I read the docs correctly they are using PP-OCRv6 which is paddleOCR",
        "posted_at": "2026-09-22T10:37:16.000Z",
        "author_url": "https://www.reddit.com/user/ephocalate/"
      }
    ],
    "evidence": [
      "[comment u/jc-atg] From our experience, Mistral OCR (not open source) is the best solution, but expensive ($2-$4/1k pages). PaddleOCR is almost as good, LightOnOCR was far behind in our testings (in most cases, Docling was better)."
    ],
    "virality": "somewhat_high",
    "post_date": "2026-09-22T03:45:45.000Z",
    "post_kind": "text",
    "post_text": "Been testing self-hosted OCR to get off Textract, since the managed services sit around $1.50 per 1,000 pages and jump hard once forms and tables are involved. The main lesson is there's no single winner, it comes down to your document type.\n\nWhat I tried:\n\ndocling: turns digital PDFs into clean markdown, great for RAG ingestion, shaky on scans and handwriting\n\nPaddleOCR-VL: multilingual VLM, 94.5% on OmniDocBench v1.5, small at ~0.9B so it fits a modest GPU\n\nMinerU: best of the bunch on formula-heavy scientific PDFs\n\nGLM-OCR: solid general image-to-markdown, good baseline to measure the rest against\n\nTwo things that caught me out: the top scores are close enough that testing on your own documents matters more than the leaderboard, and licenses bite, since Surya needs a commercial license past a revenue threshold and a couple of the leaders are CC-BY-NC.\n\nIf you'd rather not run a separate server per model, some inference servers like SIE let you swap OCR models behind one endpoint, but that's optional.",
    "sentiment": "mixed",
    "subreddit": "Rag",
    "post_flair": [
      "Tools & Resources"
    ],
    "post_title": "Best open source OCR models to replace Textract (tested a bunch)",
    "prominence": "aside",
    "source_url": "https://www.reddit.com/r/Rag/comments/1wmyoiz/best_open_source_ocr_models_to_replace_textract/",
    "entity_role": "vendor",
    "post_author": "Green-Topic-1024",
    "upvote_ratio": 0.9210526315789473,
    "mention_count": 3,
    "mention_surge": true,
    "subreddit_url": "https://www.reddit.com/r/Rag/",
    "total_upvotes": 32,
    "comments_total": 22,
    "total_comments": 22,
    "other_companies": [
      {
        "name": "PaddleOCR",
        "role": "alternative"
      },
      {
        "name": "Lighton",
        "role": "alternative",
        "domain": "lighton.ai"
      }
    ],
    "post_author_url": "https://www.reddit.com/user/Green-Topic-1024/",
    "signal_category": "feedback",
    "comments_included": 13,
    "products_mentioned": [
      "Mistral OCR"
    ]
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.