OCR. The step that turns a scan into text for a program.
OCR (optical character recognition) is the technique that recognises the letters in the image of a page, such as a scan or a photo, and turns them into text you can search, copy and have another program read. It's the first step in turning a paper document into data.
OCR gives you back the text on the page. Which number is the total and which is the customer code gets decided by a later step: data extraction.
In Italy, invoices between businesses don't need it: since 1 January 2019 they've been electronic only, and arrive with the data already structured. That leaves delivery notes, orders, documents from foreign suppliers, receipts.
Among Italian businesses using artificial intelligence, 70.8% use it to extract information from text documents, according to Istat for 2025. It's the most common use.
This entry is part of the AI and automation glossary, where every term has a short definition. Here the definition goes further: how OCR works, the difference between a native and a scanned PDF, where an Italian business needs it, and how you get from text to data.
What OCR is
OCR is optical character recognition: a program looks at the image of a page, picks out the shapes of the letters and transcribes them as text. The result is a file where you can search for a word, copy a sentence, or from which another program can read the data. Without OCR, a scan stays a photograph.
The technique has been around for decades, and in recent years artificial intelligence models have made it much more robust on crooked scans, phone photos, stamps and handwritten capitals. Joined-up handwriting, faded copies and tables with many columns and merged cells are still hard.
Native PDF or scanned PDF
A PDF can contain two different things, and only one of them needs OCR. A native PDF comes from a program, such as a document exported from your business software: the text is already in it, and you can select and copy it. A scanned PDF is an image of the page: to the eye it looks the same, but to a program it's just coloured dots.
| Document | Contents | Does it need OCR? |
|---|---|---|
| E-invoicebetween Italian businesses | Structured data, sent through the Sistema di Interscambio. |
No: the data can be read directly. |
| Native PDFexported from a program | Real text, with its layout. |
No: the text is extracted as it is. |
| Scanned PDFfrom the scanner | The image of the page. |
Yes. |
| Phone photodelivery notes, receipts | The image, often crooked or in shadow. |
Yes, after straightening and cleaning up the image. |
| Handwrittenforms, notes | The image of the handwriting. |
Yes, with mixed results: capitals hold up, joined-up writing less so. |
An easy way to tell which kind of PDF you have: try selecting a word. If only that word highlights, the text is there. If the whole page selects as a single block, you have an image, and you need OCR to read it.
Where an Italian business still needs it
Since 1 January 2019, invoices between parties resident or established in Italy can only be electronic, as reported by the Agenzia delle Entrate. In 2025, 2.4 billion of them passed through the Sistema di Interscambio, issued by 5.5 million parties. For these, OCR isn't needed: the data arrives already structured.
It's needed for everything else. Printed and signed delivery notes and transport documents. Orders that arrive as scans or photos. Invoices from foreign suppliers, for which e-invoicing is optional. And then receipts, certificates, spec sheets in PDF, contracts signed in pen and scanned back in.
From OCR to data: extraction
A business needs to know which number is the total, which is the date and which is the product code: that second step is called data extraction. Today it's often done by language models, LLMs, which read the recognised text the way a person would and fill in the fields for the business software.
It matters because retyping by hand has a small, constant error rate. In Barchard and Pace's study, about 1% of data transcribed by hand contains an error, and checking it by eye doesn't reduce that. Across a thousand values retyped, that's about ten wrong numbers, within a plausible range, and so invisible.
How to judge an OCR system
An OCR and extraction system is judged on two numbers: how many fields it extracts correctly, and what it does when it can't read something. The first is measured on a batch of the business's real documents, counting the fields one by one. The second matters more, because a total invented with confidence ends up in the business software and from there on an invoice.
The test to ask of anyone pitching you a system is simple: give it an unreadable document and see what happens. A well-built system sets it aside and flags it to a person. A badly built one fills in the fields anyway, and that's a hallucination formatted as data.
How Itria uses it
Itria's automatic document entry reads orders, delivery notes and invoices as PDFs or photos, with OCR where needed, and extracts the data in a format the business software accepts. On a test bench on 4 September 2026, with 30 documents and 237 fields, the document type was recognised in 30 cases out of 30, and 234 fields were correct, 98.7%.
The batch included an unreadable PDF, placed there deliberately, with no text in it. The system classified it as “other” and didn't invent a single field. It's a test on a test system, not at a client, and it can be rerun for anyone who asks.
Related terms
ERP
The business software that links orders, stock and invoices. It's where the data extracted from documents ends up.
LLM
The language model that reads the recognised text and decides which number goes in which field.
Unstructured data
Free text, images, emails: information a program can't use until someone puts it in order.
Data entry
The retyping by hand that OCR and extraction are there to remove.
Questions and answers
What is OCR?
OCR, optical character recognition, is the technique a program uses to look at the image of a page, such as a scan or a photo, recognise the letters and transcribe them as text.
After OCR you can search for a word in the document, copy a sentence from it, or have another program read its data. Without OCR, a scan is just a photograph.
What does a PDF with OCR mean?
It means a scanned PDF with the text recognised by OCR added underneath the image. To the eye the page looks the same, but the words can be selected and searched.
To check whether a PDF already has text, just try selecting a word: if it highlights, the text is there; if the whole page selects as a single block, it's an image.
Do you need OCR for e-invoices?
Not for Italian ones. Since 1 January 2019, invoices between parties resident or established in Italy have been electronic only, as reported by the Agenzia delle Entrate, and they arrive with the data already structured by the Sistema di Interscambio.
OCR is needed for delivery notes, orders sent as scans or photos, invoices from foreign suppliers, for which e-invoicing is optional, receipts and certificates.
How accurate is OCR?
It depends on the document: on clean print it's very accurate, on faded scans, crooked photos and joined-up handwriting much less so. The number that matters to a business is how many fields are extracted correctly from a batch of real documents, counted one by one, and what the system does with an unreadable document.
For comparison, transcription by hand gets about 1% of data wrong, according to Barchard and Pace.
What's the difference between OCR and data extraction?
OCR turns the image into text. Data extraction reads that text and decides which number is the total, which is the date and which is the product code, and puts them in the right fields.
Today extraction is often done by language models, which read the text the way a person would. A business needs both steps.
Notes on sources
- The statement on mandatory e-invoicing comes from the Agenzia delle Entrate (in Italian), given here as a faithful paraphrase and read on 26 September 2026; the same page states that e-invoicing is optional for invoices to and from abroad. The 2.4 billion invoices and the 5.5 million issuing parties are in the Agency's document I risultati 2025, pages 3 and 25.
- The 70.8% comes from Istat, Imprese e ICT, 2025 (in Italian), and is calculated on businesses with at least 10 employees that use artificial intelligence.
- The error rate of around 1% comes from Barchard and Pace, Preventing human error: The impact of data entry methods on data accuracy and statistical results, Computers in Human Behavior 27(5), 2011: an experiment on people transcribing research data, cited for the shape of the error.
- The 98.7% and the unreadable document come from an Itria measurement on a test bench on 4 September 2026: 30 documents, 237 fields, the 3 wrong ones listed one by one.
Start with the documents someone retypes today. One real batch is enough to measure what they cost.
The first step with Itria is a fifteen-minute video call: we look at which documents reach the business each week, and how many are still retyped by hand. Drop us a line about what's slowing you down. We'll make the first move: we'll look at what a customer sees when they search for you, and tell you what we found. Even if we never end up working together.