My Tool Studio
PDF Tools·3 min read

PDF to Markdown for Notes, Docs and AI Tools

Markdown has become the common language of notes apps, documentation sites and AI chat tools. PDFs, meanwhile, are where a lot of useful text is locked up. Copying from a PDF usually gives broken lines, stray page numbers and lost headings. The PDF to Markdown Converter rebuilds the structure instead, and this guide explains how it decides what is a heading, a list or a table, and how to get clean output.

PDFMARKDOWNPDF → MARKDOWN

Why copy and paste falls short

A PDF remembers positions, not meaning.

Select text in a PDF viewer and paste it, and you get every line as it was printed: paragraphs broken into short lines, words split with hyphens, headers and page numbers mixed into the middle of sentences, and tables flattened into a stream of cells. Headings look the same as body text once the font size is gone.

Markdown keeps structure in plain text. A # marks a heading, a dash marks a list item, pipes mark table cells. Rebuilding that structure is what makes the text useful again, whether a person or a program reads it next.

How structure is detected

Sizes, glyphs and alignment.

The converter reads every run of text with its font size and style. The most common size is taken as body text, and larger sizes become heading levels, largest first. Short bold lines that stand on their own become the next level down. Lines that start with a bullet or a number become list items, nested by how far they are indented.

When text lines up in columns across two or more lines, it becomes a GitHub-style table with the first row as the header. Monospaced text becomes inline code or a fenced code block that keeps its indentation. Bold, italic and links carry across as **bold**, *italic* and [text](url).

Two-column pages are read one column at a time, and paragraphs that continue into the next column are joined.

Cleaning up for AI tools

Less noise, better answers.

When you paste a document into ChatGPT, Claude or a similar tool, every repeated header and page number is noise that uses up the context window and can confuse the answer. With Remove running headers, footers and page numbers ticked, lines that repeat at the top or bottom of most pages are dropped, and lone page numbers go too.

Rejoin hyphenated words fixes words split across lines. With Between pages set to Nothing, a paragraph that runs onto the next page is joined into one. If you need to cite page numbers later, choose HTML comment instead, which adds an invisible <!-- Page N --> marker.

Pictures and scanned pages

Optional extras.

Tick Save pictures as PNG files and each picture is cut from its page, saved in an images folder, and linked from the Markdown where it appears. The download becomes a ZIP holding the .md file and the images, ready for a static site or an Obsidian vault.

Scanned pages are read with OCR when Text recognition is on Auto. OCR gives no bold or italic information, so on those pages only size-based headings, lists and paragraphs are found.

Where the Markdown ends up

One format, many destinations.

The output is CommonMark with GitHub-style tables, which is what GitHub, GitLab, Obsidian, most static site generators and many knowledge-base tools expect. Paste it into a README, a wiki page or a notes vault and the headings become an outline you can navigate.

For AI work, Markdown is a compact way to keep the document's shape. Headings tell the model where topics start, lists stay lists, and tables keep their rows and columns, which helps when you ask for a summary of one section or a comparison from a table.

A worked example

A 30 page product manual.

Take manual.pdf, 30 pages with a running header, numbered sections, spec tables and a few diagrams. Drop it in, keep Between pages on Nothing, and tick Save pictures as PNG files. The result starts with the title as a # heading, each numbered section as ##, and the spec tables as Markdown tables.

The running header and the page numbers are gone, paragraphs that crossed a page break are whole again, and the ZIP holds manual.md plus an images folder with the diagrams linked in place.

Checking the result

A two-minute review.

Heading levels are an estimate, so a quick check pays off on long documents.

  • Switch to the Preview view and scan the headings; they should match the document outline.
  • Look at tables with merged cells, which may come out as separate lines.
  • Untick Detect headings if a document with many font sizes gives too many headings.
  • Untick Turn aligned columns into tables if a form layout becomes an unwanted table.
  • Convert a page range first, such as 1-5, to test settings on a large PDF.

Try it now

Open PDF to Markdown Converter

The tool is one click away. No sign up, no upload, no payment.

Open PDF to Markdown Converter