Quick answer
Modern OCR (often called Document AI or intelligent document processing) doesn't just turn a scan into text. It understands the document: it recognizes the type, finds the fields you need, reads tables and sends the data straight into your business systems. The leap comes from vision-language models that read a page together with its layout. In practice, a person checks the result instead of retyping data.
What is OCR, and what is Document AI?
OCR (Optical Character Recognition) is technology that converts an image of text, such as a scan or a photo, into editable, searchable text. Classic OCR answers the question "which characters are on this page?"
Document AI (also IDP, Intelligent Document Processing) goes a step further and answers "what does this document say?" It knows how an invoice differs from a purchase order, recognizes that "12345" is an invoice number and that "$2,500" next to "Total Amount Due" is what you owe, and returns the data in a structure ready to save in your system.
A vision-language model (VLM) is an AI model that takes an image and text as input at the same time. For documents, this means the model "looks" at the whole page, understands its layout and can return a table, form fields or a summary directly. For more on how machines learn to see, read computer vision explained.
Why does OCR matter now?
Because paper hasn't disappeared from business. It's just wearing a digital disguise. Invoices arrive as PDFs, contracts as scans, forms as email attachments and notes as phone photos. The problem stays the same: someone has to read each document and manually move the data into your systems.
There's still plenty of room to grow. In 2025, 20% of EU enterprises with 10 or more employees used AI technologies, and the most common use was analyzing written language (11.8% of enterprises) (Eurostat, 2025).
Will e-invoicing replace invoice OCR?
Partly. Poland is a good example: its national e-invoicing system, KSeF, became mandatory for the largest taxpayers on February 1, 2026 and for everyone else on April 1, 2026 (the smallest businesses, invoicing up to PLN 10,000 gross a month, have until the end of 2026), and since February 1, 2026 all businesses must receive invoices through it (KSeF, Polish Ministry of Finance). A domestic B2B invoice arrives as a structured file, so there's nothing to read from an image.
Plenty of documents remain outside such systems, though: invoices from foreign suppliers, receipts in expense claims, contracts, handover protocols, delivery and customs documents, handwritten forms and client meeting notes. These are now the main territory for modern OCR.
How has OCR evolved from Tesseract to vision-language models?
OCR has gone through three stages: from recognizing characters, through pipelines that detect layout, to models that understand the whole page.
Classic OCR. Tesseract, one of the best-known OCR engines, was developed at Hewlett-Packard labs between 1985 and 1994. HP open-sourced it in 2005, and Google developed it from 2006 to 2017. Version 4 added an LSTM neural network engine, and today Tesseract recognizes more than 100 languages (Tesseract, GitHub). It was a big step forward with clear limits: it needed clean input, so a crumpled receipt, a skewed scan or a document with columns and tables often produced errors. It treated each page as a flat block of text with no concept of fields or layout. OCR could read the letters, but it couldn't understand the document.
Document AI pipelines. The next stage combined several models: one detects the page layout (headings, tables, fields), another reads the text, and rules or templates map values to fields. This works well for repetitive documents, but every new layout needs configuration.
Vision-language models. The breakthrough came from applying the AI techniques that had already transformed language translation and image recognition to documents. A VLM reads the page as a whole and can return Markdown, a table or JSON with fields in one pass. Costs are falling fast: the Allen Institute for AI showed that its open olmOCR model (7 billion parameters) converts a million PDF pages for about $176, while the same job through the GPT-4o API cost over $6,240 (olmOCR, arXiv 2025). Public benchmarks such as OmniDocBench now compare classic engines (including Tesseract), pipelines and VLMs on the same documents, with tables, formulas and multi-column layouts (OmniDocBench, CVPR 2025).
| Criterion | Classic OCR (e.g. Tesseract) | Document AI pipeline | Vision-language model (VLM) |
|---|---|---|---|
| Output | Flat text | Fields per template | Structure: Markdown, tables, JSON |
| Layout and tables | Weak | Good for known layouts | Good, including new layouts |
| Phone photos, skewed scans | Needs preprocessing | Depends on the pipeline | Copes better, preprocessing still helps |
| New document type | No change, it doesn't understand anyway | New configuration | Usually just a change of instructions |
| Main risk | Character errors | Brittle templates | Filling in content that isn't there |
| Hardware | CPU | CPU or GPU | GPU (on-premise or via API) |
What are the limits of vision-language models?
The biggest risk is hallucination: a model that understands context can "fill in" an illegible digit or a field that isn't on the page, and do it very convincingly. Classic OCR fails differently, returning garbage that's easy to spot. That's why VLMs need validation (do line items add up to the total, does the tax ID pass its checksum, is the date plausible), confidence thresholds and a human for uncertain cases. The second limit is hardware: VLMs need GPUs, so at high volumes you have to weigh API costs against your own infrastructure.
What does it look like in practice?
Take a typical case: a document photographed on a phone. An employee takes a picture, the system detects the page edges, corrects the perspective, reads the text and moves the required values into form fields. A person reviews the result instead of retyping the data.
Two lessons here apply broadly. First, quality starts before the model: detecting edges and straightening the image makes a big difference for documents shot on a phone. Second, the result goes into a form that a person approves, not straight into the database. That way you get the speed of AI without giving it the final word.
OCR is also the first step in building knowledge bases. In our agentic knowledge base, scans go through OCR in Polish and English, tables are converted to text with column headers, and charts are described by a vision model. Only documents prepared this way can serve as a source of answers for a language model, which we cover in how RAG makes AI smarter.
Where does Document AI help teams?
Anywhere repetitive documents slow a process down. The system extracts the right information and routes it where it's needed.
- Finance and accounting: foreign invoices, receipts in expense claims, statements, matching documents to ledger entries.
- Legal and compliance: extracting clauses and obligations from contracts, flagging unusual terms, processing regulatory documents.
- HR: resumes, onboarding paperwork, forms and personnel files.
- Customer operations: applications, forms, claims documents, less back-and-forth with customers.
- Operations and supply chain: purchase orders, shipping documents, customs paperwork, supplier files.
Reading is usually just the first step. The data still has to reach your systems and trigger the next part of the process, which is where automation comes in, as we explain in RPA vs BPA: which automation should you choose.
How do you roll out modern OCR?
Start where work piles up, and start with measurement rather than model selection.
- Pick one document type with high volume and data that's retyped by hand today.
- Measure the baseline: time per document, error count, documents per month.
- Collect a test set of real documents, including the hard ones: skewed photos, handwriting, unusual layouts. Compare solutions on it.
- Plan validation and a place for people: which fields are checked automatically, and when a result goes to manual review.
- Decide where the model runs. Documents often contain personal data, so GDPR requires knowing where they're processed. When they can't leave the company, the model runs on your own servers. We show what AI on your own infrastructure looks like on our AI infrastructure page.
- Measure after rollout and extend to the next document types.
Documents are often the hidden bottleneck in a company. The technology to remove it is available today. The question isn't whether to adopt intelligent document processing, but which document to start with.
Sources
- Tesseract OCR, repository and project history (GitHub)
- Poznanski J. et al.: olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models, arXiv 2025
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations, CVPR 2025
- Eurostat: 20% of EU enterprises use AI technologies, 2025
- KSeF: Scope of mandatory KSeF, Polish Ministry of Finance