Baidu’s AI OCR: Reading Multiple Pages Instantly — What’s The Catch?
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Baidu has open-sourced Unlimited-OCR, a large AI model that can process entire multi-page documents in a single forward pass. It introduces a new attention mechanism that maintains constant memory, enabling faster and more accurate long-document OCR. The development challenges the notion that cloud giants dominate OCR technology, emphasizing architectural improvements.

Baidu has open-sourced Unlimited-OCR, a 3-billion-parameter AI model capable of reading entire multi-page documents in a single pass, within a standard 32K context window. This development, announced on June 22, 2026, marks a significant technical achievement in OCR technology, enabling faster and more accurate long-document processing on local hardware.

The model, released under an MIT license and available on Hugging Face, builds on Baidu’s DeepSeek-OCR architecture, incorporating a new mechanism called Reference Sliding Window Attention (R-SWA). This innovation replaces the traditional linear growth in memory usage with a constant-size cache, allowing the model to process dozens of pages simultaneously without increasing latency or memory requirements. Baidu reports that Unlimited-OCR achieves a throughput of 5,580 tokens per second on OmniDocBench, outperforming its predecessor DeepSeek-OCR by approximately 12.7%.

In benchmark tests, Unlimited-OCR scores 93.23 on OmniDocBench v1.5, surpassing DeepSeek-OCR’s 87.01, with notable improvements in text edit distance and table accuracy. For long documents, such as 20- or 40-page texts, the model maintains low error rates (<0.11) across extensive reading tasks, demonstrating its suitability for real-world applications like digitizing lengthy reports or books.

Contrary to viral claims, Baidu clarifies that the model has around 8,400 downloads in the last month, not 1.9 million, and that it is less accurate than Baidu’s own PaddleOCR-VL 1.5 and Zhipu’s GLM-OCR models on certain benchmarks, though it offers advantages in multi-page, single-pass processing.

At a glance
reportWhen: announced June 22, 2026; technical repo…
The developmentBaidu announced the release of Unlimited-OCR, a 3-billion-parameter model that can parse multi-page documents instantly using a novel attention mechanism, on June 22, 2026.
Crypto market snapshot
Fear & Greed Index
25/100 — Extreme Fear
Bitcoin BTC$65,686▲ 2.5%
Ethereum ETH$1,930▲ 4.1%
Tether USDT$0.9992▲ 0.0%
BNB BNB$575.13▲ 1.9%
USDC USDC$0.9999▲ 0.0%
XRP XRP$1.13▲ 4.0%
Solana SOL$78.44▲ 3.4%
TRON TRX$0.3259▼ 0.1%
Live data · CoinGecko · alternative.me (24h change)

Implications for Local and Long-Document OCR

This development challenges the dominance of cloud-based OCR solutions by demonstrating that high-performance, multi-page document reading can be achieved locally with a single, efficient model. The architectural innovation in memory management enables faster processing of lengthy texts, reducing reliance on splitting documents into pages. For industries handling large volumes of lengthy documents—such as legal, academic, and government sectors—this could lead to more streamlined workflows and cost savings.

Furthermore, the open-source release under an MIT license makes this technology accessible for researchers and developers, potentially accelerating innovation in OCR and document understanding. However, it also raises questions about the competitive landscape, as major cloud providers like Microsoft, Google, and Azure continue to develop proprietary solutions.

Amazon

AI OCR document scanner

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Baidu’s OCR Evolution and Industry Benchmarks

Prior to this release, Baidu’s OCR efforts included PaddleOCR and DeepSeek models, which achieved strong benchmark scores but relied on traditional page-by-page processing. The recent surge in large language models and attention mechanisms has pushed the boundaries of OCR, with models like Zhipu’s GLM-OCR and others reaching high accuracy but still limited in multi-page, single-pass capabilities. Baidu’s development of R-SWA aligns with broader AI trends toward more efficient, scalable architectures that can handle complex, long-form content without linear memory growth.

The release comes amid a competitive landscape where open models are increasingly capable of challenging proprietary cloud solutions, especially in scenarios requiring local deployment, privacy, or cost efficiency.

“Unlimited-OCR demonstrates that architectural innovation in attention mechanisms can dramatically improve long-document processing speed and accuracy.”

— Baidu Research Team

Amazon

multi-page OCR software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Performance and Adoption

It is still unclear how Unlimited-OCR performs on diverse, real-world datasets outside Baidu’s internal benchmarks. While the model shows promising results in controlled tests, its effectiveness on varied document types, languages, and formats remains to be validated. Additionally, the extent to which this architecture will be adopted in commercial or enterprise settings is uncertain, as integration complexity and competing solutions may influence uptake.

Further independent evaluations are needed to confirm the model’s robustness and scalability across different hardware setups and use cases.

Amazon

long document OCR tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Development and Industry Adoption

Baidu is expected to continue refining Unlimited-OCR, potentially expanding its capabilities and benchmarks. The open-source community may develop adaptations or improvements, increasing the model’s versatility. Industry adoption will depend on real-world testing, integration ease, and comparative performance against existing cloud solutions. Major AI conferences and benchmarks in the coming months will likely provide additional insights into its practical impact.

Observers will watch whether this architecture influences other models or prompts cloud providers to innovate further in multi-page, long-document OCR.

Amazon

AI-powered text recognition device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How does Unlimited-OCR differ from previous Baidu OCR models?

Unlimited-OCR introduces a novel attention mechanism called Reference Sliding Window Attention (R-SWA), which maintains a constant memory footprint, allowing it to process multi-page documents in a single pass without increasing latency or memory usage. Previous models processed pages independently, often requiring stitching results afterward.

Can Unlimited-OCR be used on standard hardware?

Yes, the model is designed to run on standard hardware, supporting Docker, Transformers, and community quantizations, making it accessible for local deployment without specialized infrastructure.

How does the accuracy of Unlimited-OCR compare to other models?

In benchmark tests, it scores 93.23 on OmniDocBench v1.5, slightly below Baidu’s PaddleOCR-VL 1.5 and Zhipu’s GLM-OCR, which score above 94. However, its strength lies in processing entire multi-page documents in a single pass, which traditional models cannot do efficiently.

Will this technology replace cloud OCR solutions?

While it offers advantages for local processing and long documents, the extent of its adoption will depend on real-world performance, integration, and industry needs. Cloud solutions may still be preferred for their scalability and ease of use in certain contexts.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

A Brief History of AI: From Turing to Today

Merging early ideas with modern breakthroughs, this history of AI reveals a journey that has only just begun.

The Swarm Is The Weapon: Why Agentic Attacks Break The Defensive Playbook

Exploring how autonomous AI agent swarms challenge traditional cybersecurity defenses and what this means for future security strategies.