My Tool Studio
PDF Tools·4 min read

Fixing a Damaged PDF: What Can Be Saved and How

The PDF opened fine last week. Now your reader says the file is damaged and can't be repaired, or it opens to blank pages, or it won't open at all. Sometimes the file really is broken, sometimes only its index is, and sometimes it was never a PDF in the first place. This guide explains what goes wrong inside a PDF, how the Repair PDF tool tries to recover it, and how to tell a file that can be fixed from one that needs to be downloaded again.

INOUTIN → OUT

How a PDF breaks

The index is fragile. The content often isn't.

A PDF is a collection of numbered objects, such as pages, fonts, images and text streams, followed by a cross-reference table that says where each object starts in the file. At the very end is a short trailer pointing to that table and an end-of-file marker. A reader jumps to the end first, finds the table, and uses it to find everything else.

That design makes the end of the file the weakest point. If a download stops early, the table and trailer are missing, and many readers give up even though most of the pages arrived. If an app adds a few bytes at the start, every offset in the table is wrong. In both cases the page content is still there and a repair tool can often find it by scanning the whole file for objects.

What the diagnosis tells you

Read this before you download anything.

Repair PDF scans the file before it tries to fix anything, and lists what it finds in plain English. A missing end-of-file marker usually means a cut-off download or copy. Stray bytes before the header point to an email gateway or upload system that added data. Extra bytes after the end come from padding or a failed save.

The most useful finding is often that the file isn't a PDF. Web pages saved from a broken download link, Word files with the wrong extension, and photos renamed by mistake all show up here. The tool offers the file back with the right extension, and turns JPG and PNG images into a PDF page.

The four recovery methods

Each one catches what the others miss.

Every method runs in turn, and its output is opened again and checked page by page:

  • qpdf structure rebuild. The qpdf engine, running in the browser as WebAssembly, rebuilds the cross-reference table from the objects it finds and rewrites the file. It keeps bookmarks, forms and encryption.
  • pdf-lib object reload. Every object is read one by one, the pages are copied into a fresh file, and a file cut off mid-object is read up to the last complete one.
  • Page tree rebuild. When pages exist but the list that links them is broken, loose page objects are collected and linked again, in object order.
  • Image rescue. Pages that can still be drawn are saved as pictures. The look survives, the selectable text doesn't.

A worked example: a half-downloaded statement

Some pages arrived. Save those.

Say statement-march.pdf was meant to be 24 pages, but the download stalled and your reader refuses to open it. Drop it on the tool and click Repair PDF. The diagnosis says the end-of-file marker and the cross-reference pointer are both missing, which confirms a cut-off download.

The methods table shows the result. In this case qpdf can't rebuild the file, because too much of the end is missing. pdf-lib reads up to the last complete object and recovers 18 pages. The tool marks the pdf-lib result as recommended, and the verdict reads partly recovered, 18 of about 24 pages.

Preview the 18 pages to see what you have. For many tasks, such as finding one transaction, that's enough. For a complete statement, the missing six pages can only come from a fresh download, because they were never saved to your device.

When to use image rescue

The last resort, but a good one.

Sometimes a page's objects are damaged in ways that break copying, yet the PDF viewer can still draw the page. Image rescue captures exactly that: every page the viewer can draw is saved as a 150 DPI image in a new PDF. By default it runs only when the other methods recover fewer pages than the viewer can see. Set it to Always to get it as an extra choice every time.

Rescued pages look right but their text is gone, so run the result through OCR PDF if you need to search or copy from it.

Habits that prevent damaged PDFs

Cheaper than any repair.

Most broken PDFs come from a handful of situations, and all of them are avoidable:

  • Let downloads finish, and check the file size against what the site shows.
  • Eject USB drives and memory cards before unplugging them.
  • Don't edit a PDF that is still syncing to a cloud folder.
  • Keep a second copy of anything you can't easily get again.
  • When a download fails, download it again before trying a repair.

Try it now

Open Repair PDF

The tool is one click away. No sign up, no upload, no payment.

Open Repair PDF