My Tool Studio
PDF

PDF to Text Converter

This PDF to Text converter pulls the text out of a PDF as a plain .txt file. Choose Reading order to join lines into paragraphs, Line by line to keep each printed line, or Keep layout to line up columns and tables with spaces, like a printout. Two-column pages are read one column at a time. Scanned pages have no text layer, so they are read with OCR in 29 languages. Running headers, footers and page numbers can be removed, words split with a hyphen are rejoined, and you can mark page breaks with a label, blank lines or a form feed. See word and character counts, copy the text, download one .txt, or get one file per page in a ZIP. Several PDFs can be done at once.

Always freeNo sign upRuns in your browser

Leave empty for all pages, or type pages and ranges like 1-3, 5, 8-.

Scanned pages have no text layer. OCR reads them from the page image.

The first run downloads this language's data.

For pages that mix two languages, such as Hindi and English.

Good to know

  • Text inside pictures is only read when OCR runs on that page.
  • Simple two-column pages are read column by column. Magazine layouts with boxes and captions can come out in a mixed order.
  • Keep layout lines up columns with spaces, so view it in a monospaced font such as Notepad or VS Code.
  • Password-protected PDFs must be unlocked first.

How to use

01

Add PDFs

Drop one or more PDFs of up to 200 MB each. Type a page range like 1-5 if you only need part of the document.

02

Pick layout and cleanup

Choose Reading order, Line by line or Keep layout, and what goes between pages. Set OCR for scans, and tick header removal, hyphen repair or Windows line endings.

03

Extract and save

Press Extract text. Check the counts, copy the text, press Download .txt, or choose One .txt per page for a ZIP.

Why PDF to Text Converter

Choosing a text layout

Each layout suits a different job.

LayoutWhat you getGood for
Reading orderParagraphs joined, blank line between blocksReading, AI tools, word counts
Line by lineOne line per printed line, tabs between columnsAddresses, lists, data cleanup
Keep layoutText placed by position with spacesTables and forms in a monospaced viewer

Text layer or OCR

Most PDFs made from Word, a browser or accounting software have a text layer, and extracting it is instant and exact. Scans and photos of pages do not, so OCR has to recognise the letters from the page image. OCR is slower, needs the right language, and can misread small or blurred text, so check numbers and names.

Common questions

How do I extract all the text from a PDF?
Drop the PDF and press Extract text. The text appears in a box where you can copy it, and Download .txt saves it as a UTF-8 text file.
What is the difference between the three PDF to text layouts?
Reading order joins lines into paragraphs and suits reading or pasting into other tools. Line by line keeps each printed line. Keep layout places text by its position with spaces, so columns and simple tables stay lined up when viewed in a monospaced font.
Can I get text from a scanned PDF with this converter?
Yes. Keep Text recognition on Auto and pick the language of the scan. Pages with no text layer are read with OCR. Set it to Always if a page has a broken text layer that gives gibberish.
Why does text from a two-column PDF come out in the right order?
Each page is checked for a clear vertical gap between columns. When one is found, the left column is read before the right, and lines that span both columns, such as a title, keep their place.
How do I remove page numbers and headers from extracted PDF text?
Keep Remove running headers, footers and page numbers ticked. Lines in the top and bottom bands that repeat on at least half of the pages, and lone page numbers, are dropped.
Can I save each PDF page as its own text file?
Yes. After extracting, press One .txt per page (ZIP). Each page is saved as page-001.txt, page-002.txt and so on, in a folder per PDF when you converted several.
Why does the extracted text show strange symbols instead of words?
Some PDFs use fonts that map letters to private codes, so the text layer is not real text. The converter spots most of these and uses OCR in Auto mode. If a page still shows symbols, set Text recognition to Always.
What does the form feed page separator do?
It puts a form feed character between pages, the same marker the pdftotext command uses. Many text tools and printers treat it as a page break, and it is easy to split on in scripts.
Does PDF to Text upload my file?
No. The PDF is read with pdf.js in your browser, and OCR runs with Tesseract.js on your device. Only the OCR language data is downloaded the first time.

More PDF tools

View all