AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Baidu released Unlimited-OCR, a 3-billion-parameter AI model capable of processing entire multi-page PDFs in one pass. This innovation improves speed and memory efficiency, impacting OCR technology for long documents.

Baidu has open-sourced Unlimited-OCR on June 22, 2026, a new AI model designed to parse entire multi-page documents in a single forward pass. This development marks a significant step in AI-driven OCR technology, especially for processing long PDFs efficiently. The model is available under an MIT license, supporting multiple deployment frameworks, and is positioned as a major technical achievement in document parsing.

The Unlimited-OCR model is built on Baidu’s DeepSeek-OCR architecture, incorporating a novel Reference Sliding Window Attention (R-SWA) mechanism. This innovation replaces the traditional linear growth of memory with a constant-size cache, enabling the processing of dozens of pages in one pass without increasing latency or GPU memory usage. Baidu reports that this results in a throughput of approximately 5,580 tokens per second, with a 12.7% performance increase over previous models like DeepSeek-OCR.

According to the technical report published on arXiv, on the OmniDocBench benchmark, Unlimited-OCR scores 93.23 overall, outperforming its predecessor DeepSeek-OCR, which scored 87.01. It demonstrates high accuracy on long documents, maintaining a low error rate even for 40+ page texts, with an edit distance of 0.1069. However, it is not the highest-scoring open OCR model overall, with PaddleOCR-VL 1.5 and Zhipu’s GLM-OCR scoring slightly higher on some benchmarks.

Contrary to viral claims, Baidu’s model has approximately 8,400 downloads in the last month, not 1.9 million, indicating high but not viral popularity. The model is available with community support for various frameworks, including Hugging Face, Docker, and llama.cpp, making it accessible for local deployment.

At a glance
reportWhen: announced June 2026
The developmentBaidu launched Unlimited-OCR, an open-source AI model that processes multi-page documents in a single forward pass, enhancing long-document OCR.

Implications for Long-Document OCR and AI Efficiency

This development offers a practical solution for processing lengthy PDFs and complex documents without splitting or page-by-page OCR, which often causes errors in reading order and cross-references. Its architecture allows faster, more memory-efficient processing, potentially transforming industries reliant on large-scale document digitization, legal work, and academic research. While it does not surpass the top benchmark scores of some models, its ability to handle long texts in a single pass addresses a key limitation of existing OCR systems, making it highly relevant for real-world applications.

Amazon

PDF OCR software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Baidu’s OCR Model Lineage and Industry Landscape

The Unlimited-OCR model is a refinement of Baidu’s earlier DeepSeek-OCR, which itself was influenced by open models like PaddleOCR and PaddleOCR-VL. Baidu’s focus on architectural improvements—particularly the R-SWA mechanism—addresses longstanding issues of memory growth and latency in decoder-based OCR models. This approach aligns with broader industry efforts to improve long-document processing, where splitting pages often compromises accuracy. Prior to this release, models like PaddleOCR-VL and Zhipu’s GLM-OCR led in benchmark scores, but lacked the ability to process entire documents in one pass.

Open-source AI models for OCR have become increasingly prominent, with Baidu’s release positioning itself as a practical, deployable solution that balances accuracy with efficiency, especially for lengthy texts. The release also counters viral narratives suggesting China’s OCR efforts are obsolete, emphasizing that Baidu’s innovations are architectural rather than purely benchmark-driven.

“Unlimited-OCR demonstrates that constant memory and single-pass processing are achievable for multi-page documents, significantly improving speed and efficiency.”

— Baidu Research Team

Remaining Questions About Model Performance and Adoption

It is not yet clear how Unlimited-OCR performs on independent, external long-document benchmarks beyond Baidu’s internal tests. Its accuracy relative to top benchmark scores is slightly lower, and real-world deployment challenges—such as handling complex layouts or multi-language documents—remain untested at scale. The actual adoption rate outside Baidu’s ecosystem is also uncertain, as the model’s popularity appears limited to around 8,400 downloads in recent weeks.

Future Developments and Industry Impact of Baidu’s OCR Breakthrough

Baidu is expected to continue refining Unlimited-OCR, possibly integrating it into commercial products or APIs. Industry observers anticipate that other OCR providers may adopt similar memory-efficient architectures, pushing the field toward more scalable long-document processing solutions. Further independent testing and real-world case studies will clarify its practical advantages and limitations, shaping future OCR research and deployment strategies.

Key Questions

What makes Baidu’s Unlimited-OCR different from previous models?

It introduces a Reference Sliding Window Attention (R-SWA) mechanism that keeps memory usage constant regardless of document length, enabling processing of entire multi-page PDFs in a single pass.

How does the performance of Unlimited-OCR compare to other OCR models?

On benchmark tests like OmniDocBench, it scores slightly lower than top models like PaddleOCR-VL 1.5 and Zhipu’s GLM-OCR but offers superior long-document handling due to its architectural design, with a focus on efficiency rather than peak accuracy.

Can I run Unlimited-OCR locally?

Yes, the model is available under an MIT license with support for frameworks like Transformers, Docker, and llama.cpp, making local deployment feasible for developers and organizations.

Will this model change how long documents are processed in industry?

Potentially, yes. Its ability to process entire documents in one pass can improve accuracy and speed in legal, academic, and enterprise workflows, reducing errors caused by splitting pages or multiple passes.

Source: ThorstenMeyerAI.com

You May Also Like

Disk Is the Contract: Inside Threlmark’s Local-First Architecture

Threlmark’s innovative approach uses on-disk JSON files as the single source of truth, enabling open, portable, and restartable project management without a database.

Examining The Reliability Of AI Sovereignty Certifications Through The 24% Rule

An analysis of the reliability of AI sovereignty certifications, focusing on the 24% ownership rule and its implications for data control and legal jurisdiction.

Game 3: Both Teams Destroy Inhibitors?

In Game 3, both teams destroyed inhibitors simultaneously, marking a rare strategic move. The event impacts team strategies and playoff dynamics.

Unlocking the PSP’s Dual Core Setup

A recent development reveals how to unlock the PSP’s dual-core setup, enabling enhanced performance. Here’s what is confirmed and what remains unclear.