ToolXkit Icon

Technology · 6 min read · Updated 3 August 2026

How OCR works, and how to get accurate results from it

Optical character recognition is not magic and it is not guessing. Understanding the pipeline is what lets you fix a bad result.

Optical character recognition turns a picture of writing into text a computer can search, copy and edit. When it works, it feels like magic. When it does not, the output is a garbled mess and the reason is not obvious. The problem is almost always in the input, and once you understand the steps OCR goes through, you can fix it.

What actually happens

Every OCR engine runs roughly the same five stages.

1. Preprocessing

The image is cleaned up. It is converted to greyscale, then thresholded so every pixel becomes either ink or paper. Skew is corrected: a page scanned two degrees crooked is straightened, because every later step assumes text runs horizontally. Speckle from a dirty scanner bed is removed.

This stage does more for accuracy than any other. A clean, straight, high-contrast image wins most of the battle. Tesseract's own guide to improving output quality covers these same steps: rescaling, binarisation, noise removal and deskewing.

2. Layout analysis

The page is divided into regions: blocks of text, images, tables, rules. Then each text block is split into lines, and each line into words. This is where documents with several columns go wrong. If the engine reads across the columns instead of down them, every word is correct and the sentences are nonsense.

3. Character recognition

Each character shape is classified. Older engines matched shapes against templates. Modern ones use a neural network trained on millions of examples, and they usually read a whole line at a time instead of single letters. That way the shapes of nearby characters help with each decision. (Tesseract 4 and later work like this, with an LSTM line recogniser.)

4. Language modelling

People forget this stage, and it is why choosing the right language makes such a difference. The engine holds a dictionary and a model of which letter sequences are likely. When a shape is ambiguous, it uses context: rn versus m, 1 versus l versus I, 0 versus O. Set the wrong language and you lose this correction completely. The engine will confidently produce English-shaped nonsense from a German page.

5. Output

The recognised text comes out either as plain text or written back into the PDF as an invisible layer on top of the original scan. The second option is usually what you want: the page still looks exactly like the scan, but Find works and you can select text.

What ruins accuracy

Problem Effect Fix
Low resolution Severe. Below 200dpi accuracy falls off a cliff Rescan at 300dpi
Skew or curl Lines get misread or merged Flatten the page; scan, do not photograph
Low contrast Faint characters dropped completely Increase contrast before OCR
Wrong language set Plausible words, wrong ones Set the actual language
Decorative or script fonts Often unrecoverable Nothing reliable; expect to retype
Handwriting Standard OCR is not built for it Needs a handwriting-specific model
Shadow across the page Half the page thresholds to solid black Even lighting; use a scanner app’s document mode

Getting the best result

  1. Scan at 300dpi. Not 150, which is too coarse for small type. Not 1200, which is slow and adds nothing.
  2. Scan in greyscale, not colour, unless colour carries meaning. It is faster and thresholds more cleanly.
  3. Get the page flat and square. A book pressed properly onto the glass beats any amount of software correction.
  4. Taking a photo instead? Use even light and shoot straight down. Our Scan to PDF tool has a document mode that turns the background white, which is exactly the preprocessing OCR wants.
  5. Set the correct language. It is the cheapest way to gain accuracy.
  6. Check the numbers by hand. Digits have no dictionary to fall back on, so errors hide there. If the document contains figures that matter, proofread them.

Running OCR without uploading anything

OCR has usually meant either desktop software or a cloud service. The cloud option means handing over whatever you are scanning, which rules it out for medical records, contracts and anything covered by a confidentiality agreement.

Our OCR PDF tool runs the recognition engine as WebAssembly inside your own browser. The engine downloads once, around 12MB, and then works offline. Nothing about your document is sent anywhere, because there is nowhere for it to be sent to.

It is slower than a GPU in a datacentre, since it really is reading every page on your laptop. For a two-hundred-page scan, expect to leave it running. That is the trade-off, and for anything confidential it is an easy one to accept.

After OCR

Once a scan is searchable, it behaves like any other text-based PDF. You can extract the text, convert it to Word, or search it in place. Run OCR first and every other tool on this site starts working on a document that used to be just pictures.

Tools mentioned in this article

Keep reading