My Tool Studio
PDF Tools·4 min read

How to Make a Scanned PDF Searchable with OCR

You scanned a stack of signed contracts last year, and today you need the one that mentions a specific clause. Ctrl+F finds nothing, because every page is a photo of paper. OCR fixes that by reading the words in each image and storing them in the PDF as real text. This guide walks through doing it with the OCR PDF tool, explains what changes inside the file, and covers the settings and scanning habits that decide whether the result is useful or full of mistakes.

INOUTIN → OUT

Why a scan can't be searched

A picture of words is not words.

When a scanner or a phone app makes a PDF, it takes a photo of each sheet and wraps the photos in a PDF file. Your PDF reader shows the photo, and to your eye it looks like any other document. To the computer, though, each page is one large image with no letters in it. Search, copy and paste, and screen readers all need characters, and there are none.

Some scan apps run OCR for you, and some don't. Office copiers often save plain image PDFs by default. A quick test tells you which kind you have: try to highlight a line of text. If the whole page highlights as a block, or nothing happens, the file needs OCR.

What OCR adds to the file

Invisible text, placed word by word.

Optical character recognition looks at the image, finds shapes that match letters, and returns each word with its position on the page. The OCR PDF tool uses Tesseract, a long-running open source engine, compiled to run inside your browser. It draws each page at the resolution you choose, reads it, and gets back every word with its box.

Those words are then written into the PDF as invisible text, each one laid over the matching word in the picture. The image is not changed at all. That's why the result looks identical to the scan, yet when you drag across a line, the highlight follows the printed words, and when you search for a name, the reader jumps to the right spot.

The invisible layer uses a special font that can hold any script. A Hindi or Arabic page copies out as real Hindi or Arabic text, not as boxes or question marks.

A worked example: a 12-page lease

From photo pages to a searchable file.

Say you have lease-2024.pdf, twelve pages scanned on an office copier, with the first page in English and a Gujarati addendum on pages 11 and 12. Drop the file on the tool. Set Document language to English and Second language to Gujarati, so both scripts are read on every page. Leave Pages blank to cover the whole lease, and keep Accuracy on Standard.

Click Make PDF searchable. The first run downloads the OCR engine and the two language files, which your browser then keeps for next time. Progress moves page by page. When it finishes, the results show how many pages were read, how many words were added and an average confidence score. The page-by-page report shows the confidence for each page, and the recognised text appears in a box underneath.

Download lease-2024-ocr.pdf and open it in your usual reader. Search for the tenant's name and it should jump straight to page 1. Search for a Gujarati word from the addendum and it should land on page 11. If one page has a much lower score than the rest, look at it: a crooked or faint page is the usual reason.

Choosing languages and accuracy

Two settings do most of the work.

Language matters more than anything else. Tesseract uses the language to decide which shapes are letters, so reading a Tamil page as English doesn't give a rough result, it gives nonsense. Pick the main language of the document, and add a second one only when pages really mix scripts, since each extra language slows reading down a little.

Accuracy sets how large each page is drawn before it's read. Fast works at 150 DPI and suits large, clean print. Standard, at 200 DPI, handles most office scans. High reads at 300 DPI and helps with footnotes, dense tables and small print, at the cost of time. Start with Standard, and only move up if the confidence score comes back low.

  • Skip pages that already have selectable text stays on by default, so mixed files only OCR the scanned pages.
  • Improve contrast helps faint photocopies, grey pencil and yellowed paper.
  • Rebuild pages from the scan replaces a garbled old text layer with fresh OCR text.

Mistakes that make OCR look broken

Most bad results start at the scanner.

When OCR output is poor, the cause is usually one of these:

  • The wrong language. Always the first thing to check when a page comes out as random letters.
  • Tilted pages. A page scanned at an angle reads far worse than one that's a little blurry.
  • Very low resolution. Phone photos taken from far away, or scans at 100 DPI, lose the detail OCR needs.
  • Handwriting. Tesseract is built for print, so notes in the margin are mostly missed.
  • Expecting perfection. Even a good scan can confuse 0 and O or 1 and l, so check account numbers and amounts by eye.

Keeping private scans private

The file never leaves the tab.

Scanned documents are often the most sensitive ones people own: signed contracts, bank statements, medical letters, ID cards. With this tool the PDF is read and rewritten inside your browser, and nothing about the file is uploaded. The only downloads are the OCR engine and language data, which contain nothing about your document.

Once a scan is searchable, other tools get more useful too. You can search a long archive in seconds, copy a clause into an email, compare two versions of an agreement with Compare PDF, or find and remove a name with Redact PDF.

Try it now

Open OCR PDF

The tool is one click away. No sign up, no upload, no payment.

Open OCR PDF