My Tool Studio
PDF

OCR PDF

OCR PDF turns a scanned PDF into a searchable one, right in your browser. Scanners and phone scan apps save each page as a picture, so you can't search, select or copy a single word. This tool reads every page with Tesseract OCR and lays an invisible text layer over the original image. The page looks exactly the same, but Ctrl+F, copy and paste, and screen readers now work. Pick from 44 languages, add a second language for mixed documents such as Hindi and English, and choose which pages to read. Pages that already have text are skipped. You also get the recognised text as a .txt file. Nothing is uploaded, so it suits contracts, bank statements and scanned IDs.

Always freeNo sign upRuns in your browser

How to use

01

Add your scanned PDFs

Drop one or more PDFs on the upload area, up to 200 MB each. Each file shows its page count once it has been read, so you can check you picked the right scan.

02

Choose the language and pages

Pick the document language, and a second one if the text mixes scripts. Leave Pages blank for the whole file or type ranges like 1-3, 7. Standard accuracy suits most office scans, and High reads small print at 300 DPI.

03

Run OCR and download

Click Make PDF searchable and watch the page-by-page progress. When it finishes, download the searchable PDF, save the recognised text as a .txt file, or copy it straight to your clipboard.

Why OCR PDF

What happens to each page during OCR

A scanned page is just a photo of paper. To make it searchable, the tool draws the page at the resolution you chose, reads it with the Tesseract engine, and gets back every word with its position on the page. Those words are then written into the PDF in invisible ink, one word over each word in the picture.

Because the text is invisible and sits in the right place, highlighting a line in your PDF reader selects the words you see, and searching for a name jumps to the spot where it appears. Screen readers can read the page aloud, and the text can be copied into an email or a spreadsheet.

  • The text layer uses a special font that covers every script, so Hindi, Arabic and Chinese text copies correctly.
  • Pages you leave out of the range are copied into the new file unchanged.
  • Links, bookmarks and form fields on the original pages stay in place with Keep the original pages.

Choosing the accuracy setting

Higher settings draw the page larger before reading it, which helps with small print but takes longer. Start with Standard and only move up if the confidence score is low.

SettingResolutionGood for
Fast150 DPILarge, clear print such as typed letters and slides
Standard200 DPIMost office scans, invoices and statements
High300 DPISmall print, footnotes, dense tables and older copies

Getting cleaner OCR from a poor scan

OCR can only read what the image shows. A few habits at scan time make a bigger difference than any setting afterwards.

  • Scan at 300 DPI in grayscale when you can. Colour adds size without helping OCR.
  • Keep pages straight. A tilted page reads worse than a slightly blurry one.
  • Tick Improve contrast for faint photocopies, pencil marks or yellowed paper.
  • Choose the right language. Reading a Gujarati page as English gives nonsense, not a partial result.

Common questions

What does OCR PDF actually change in my scanned file?
It adds text. Each page is read with Tesseract OCR and the words are written into the PDF as an invisible layer, placed exactly over the matching words in the image. The picture of the page is not touched, so the file looks the same but can now be searched, selected and copied.
Will the searchable PDF look different from my original scan?
Not with the default Keep the original pages option, because the scan stays as it is and only invisible text is added. The Rebuild pages option replaces each chosen page with a fresh image at the accuracy setting you picked, which can change the sharpness and the file size a little.
Which languages can the OCR PDF tool read?
44 languages, including English, Hindi, Gujarati, Marathi, Bengali, Tamil, Telugu, Urdu, Arabic, Spanish, French, German, Russian, Chinese, Japanese and Korean. You can add a second language for pages that mix two scripts. Each language file downloads the first time you use it and your browser keeps it for next time.
How accurate is OCR on a scanned PDF?
It depends mostly on the scan. Clean, straight, printed pages usually read very well, while faint copies, skewed photos, tiny print and handwriting read poorly. The results show an average confidence per page. If it is low, check the language, switch to High accuracy, or tick Improve contrast and run it again.
Why did the OCR tool skip some of my pages?
Those pages already had selectable text, so OCR would only add a second copy of it. A page counts as having text when it holds a few dozen characters or more. Untick Skip pages that already have selectable text to force OCR on every page, which helps when the existing text is wrong or garbled.
Is it safe to OCR confidential PDFs with this tool?
Yes. The PDF is read and rewritten by code running in your browser tab, and it is never sent to a server. The only downloads are the OCR engine and the language data, which come from a public code library and contain no information about your file.
How long does OCR take for a long PDF?
Expect several seconds per page on a laptop, longer on a phone or at High accuracy. The very first run also has to download the OCR engine and language data. Progress is shown page by page, and you can press Stop at any time. Pages finished before you stop are not saved, so for very long files OCR a range at a time.
Can I OCR a password-protected PDF?
Not directly. The tool needs to rewrite the file, and an encrypted PDF can't be rewritten without its password. Remove the password with the Unlock PDF tool first, run OCR on the unlocked copy, then add a password again with Protect PDF if you need one.
When should I use Rebuild pages instead of Keep the original pages?
Use Rebuild pages when a PDF already has text but copying it gives nonsense, such as boxes, random letters or missing spaces. That happens with old fonts that don't map to real characters. Rebuilding replaces the page with its image plus fresh OCR text, which removes the broken layer. For normal scans, keep the original pages.
Can I get just the text from a scanned PDF?
Yes. After OCR the recognised text appears in a box under the results, marked with page numbers. Click Copy text to paste it elsewhere, or Download text to save a .txt file. The searchable PDF is still there if you want it as well.

More PDF tools

View all