My Tool Studio
PDF Tools·3 min read

PDF to Text: Clean .txt From Any PDF

Plain text is still the most useful format for a lot of jobs: searching, counting words, feeding a script, pasting into a form or an AI tool. Getting it out of a PDF should be easy, yet copy and paste often scrambles columns, splits words with hyphens and drags in page numbers. The PDF to Text Converter gives you control over how the text is laid out, handles scanned pages with OCR, and runs entirely in your browser. This guide explains each mode and option.

PDFTEXTPDF → TEXT

Three layouts for three jobs

The same text, shaped differently.

Reading order joins the lines of each paragraph into one, keeps headings and list items on their own lines, and separates blocks with a blank line. It is the version to read, count or paste into another tool.

Line by line keeps each printed line as a line, in reading order, with a tab where the page had a wide gap. That suits address lists, product codes and data cleanup, where every line is a record. Keep layout places text by its position, padding with spaces so columns and simple tables line up. Open it in a monospaced font such as Notepad or VS Code to see the effect.

Columns and reading order

Left column first.

A two-column page read straight across mixes the two columns line by line. The converter looks for a clear vertical gap down the middle of each page. When it finds one, it reads the left column, then the right, and keeps lines that span both columns, such as a title, in their place. Pages with boxes, captions and sidebars can still come out in a mixed order, because there is no single reading path to recover.

Scanned PDFs and OCR

When there is no text to extract.

A scanned PDF holds pictures of pages. With Text recognition on Auto, pages without a usable text layer are rendered and read with Tesseract OCR in the language you choose, from English to Hindi, Gujarati, Bengali, Tamil, Urdu and Arabic. The statistics line shows how many pages needed OCR.

Some PDFs have a text layer made of private symbol codes, which extracts as gibberish. Most are detected and sent to OCR automatically. If a page still looks wrong, set Text recognition to Always.

Text layer or OCR: how to tell

Try selecting a word.

Open the PDF in any viewer and try to select a word. If it highlights, the page has a text layer and extraction will be instant and exact. If the whole page highlights as one picture, it is a scan, and OCR will be needed. Mixed PDFs are common, such as a typed report with a few scanned signature pages, and Auto mode handles each page on its own.

OCR speed depends on your device. A modern laptop reads a page in a few seconds; an older phone takes longer. The Stop button ends the job early if you picked the wrong language.

A worked example

A 40 page thesis.

Take thesis.pdf, 40 pages with a running title at the top, page numbers at the bottom and a two-column literature review. Keep Reading order, tick header removal and hyphen repair, and set Between pages to Nothing. The result reads as one continuous text: the running title and page numbers are gone, the two-column section is read column by column, and the word count above the text gives the total for the body.

Cleanup options

Removing what you did not want.

Remove running headers, footers and page numbers drops lines that repeat at the top or bottom of at least half of the pages, plus lone page numbers. Rejoin words split with a hyphen at line ends fixes words broken across lines, but only when the next line starts with a lowercase letter, so real hyphenated names are left alone.

Between pages controls what separates pages: a --- Page N --- label, blank lines, a form feed character like the pdftotext command uses, or nothing. Windows line endings writes CRLF for older Windows tools.

Saving the text

One file, many files, or the clipboard.

Above the text, a line shows the page count, words and characters, which is a quick check that nothing went missing. Copy puts the text on the clipboard, and Download .txt saves a UTF-8 file named after the PDF. One .txt per page saves page-001.txt, page-002.txt and so on in a ZIP, which is useful for splitting a long document for review or processing. With several PDFs, each gets its own file, or its own folder of pages.

  • Word count for an essay: Reading order, headers removed.
  • Importing a price list into a spreadsheet: Line by line, then paste into Excel.
  • Checking a table's alignment: Keep layout in a monospaced editor.
  • Feeding a long report to a script: one .txt per page.

Try it now

Open PDF to Text Converter

The tool is one click away. No sign up, no upload, no payment.

Open PDF to Text Converter