Baidu Releases MIT-Licensed 3B OCR Model for long documents.
Article excerpt
Highlighted: the sentence this signal was extracted from
Baidu Releases MIT-Licensed 3B OCR Model for long documents. Tl;dr. * Baidu's Unlimited-OCR is a 3-billion-parameter MIT-licensed model that processes multi-page PDFs in a single inference pass. * The model has a 32,768-token context window and supports vLLM, SGLang, Ollama, llama.cpp, and Hugging Face Transformers. * No benchmark results are included in the release; training data and language coverage are also not disclosed. Baidu published Unlimited-OCR to Hugging Face, a 3-billion-parameter model for document parsing released under an MIT license. The headline capability is what the model card calls "One-shot Long-horizon Parsing": processing multi-page PDFs and image stacks in a single inference pass rather than requiring documents to be pre-sliced page by page. A 32,768-token context window supports that approach on longer documents. Most open-source OCR pipelines force you to cut input into individual pages, run each through a model separately, then stitch the outputs back together. The architectural bet here is that a sufficiently long context window lets the model handle that coherence itself. According to the model card, Unlimited-OCR builds on DeepSeek-OCR and DeepSeek-OCR-2, and deploys via Hugging Face Transformers, vLLM, SGLang, Docker, Ollama, and llama.cpp, meaning it should slot into most existing inference setups without significant rework. The MIT...
Keep reading with a free account
The rest of this article, and every signal for Baidu, is in your free account.
