Drop files here, or click to browse
PDF files up to 200 MB each · several files allowed
Leave empty for all pages, or type pages and ranges like 1-3, 5, 8-.
Scanned pages have no text layer. OCR reads them from the page image.
The first run downloads this language's data.
For pages that mix two languages, such as Hindi and English.
Good to know
- Text inside pictures is only read when OCR runs on that page.
- Simple two-column pages are read column by column. Magazine layouts with boxes and captions can come out in a mixed order.
- Keep layout lines up columns with spaces, so view it in a monospaced font such as Notepad or VS Code.
- Password-protected PDFs must be unlocked first.
How to use
Add PDFs
Drop one or more PDFs of up to 200 MB each. Type a page range like 1-5 if you only need part of the document.
Pick layout and cleanup
Choose Reading order, Line by line or Keep layout, and what goes between pages. Set OCR for scans, and tick header removal, hyphen repair or Windows line endings.
Extract and save
Press Extract text. Check the counts, copy the text, press Download .txt, or choose One .txt per page for a ZIP.
Why PDF to Text Converter
- Three text layouts: paragraphs, printed lines, or columns kept with spaces.
- OCR for scanned pages in 29 languages, two at a time for mixed pages.
- Removes running headers and page numbers and rejoins hyphenated words.
- Batch extraction, one file per page, and no upload.
Choosing a text layout
Each layout suits a different job.
| Layout | What you get | Good for |
|---|---|---|
| Reading order | Paragraphs joined, blank line between blocks | Reading, AI tools, word counts |
| Line by line | One line per printed line, tabs between columns | Addresses, lists, data cleanup |
| Keep layout | Text placed by position with spaces | Tables and forms in a monospaced viewer |
Text layer or OCR
Most PDFs made from Word, a browser or accounting software have a text layer, and extracting it is instant and exact. Scans and photos of pages do not, so OCR has to recognise the letters from the page image. OCR is slower, needs the right language, and can misread small or blurred text, so check numbers and names.