Skip to content

Glossary

What is OCR?

OCR, or optical character recognition, is technology that turns text in images and scanned documents into machine-readable text.

  • Also called Optical character recognition

What it means

OCR reads the letters and numbers in a picture of text, such as a scanned invoice, a photographed receipt or a PDF made from images, and outputs text a computer can search and process.

Modern OCR models also keep layout information, such as where each paragraph or table cell sits. Combined with a language model, OCR output can be turned into structured fields like supplier, date and total.

Why it matters for a business

Paper and scanned documents still carry a lot of business data. OCR removes the retyping, and with confidence scores it can flag unclear documents for a person to check.

A business example

Things to watch

  • Accuracy drops with blurry photos and handwriting.

  • Tables and multi-column layouts need testing.

  • Use confidence scores to route unclear pages to people.

  • Handle personal data in documents carefully.

Keep exploring

FAQ

Questions about OCR

Have a question that is not here? Ask us directly.

Start a project

Tell us what you want to build. We will show you a faster path.

Send a short brief. We reply with questions, a suggested plan and an estimate you can compare with other offers.

Your privacy choices

We use necessary storage to run this site. With your permission we also use Google Analytics to see which pages help people, and load maps from Google. You can change this at any time. Read the cookie policy.