My Tool Studio
Text Tools·4 min read

Extracting Links From Documents for a Link Audit

A link audit starts with a list, and building the list is the boring half of the job. Doing it by hand means scrolling, selecting, and inevitably missing the one address that mattered. The URL Extractor does it in a paste: drop in a report, a draft post, an exported newsletter or raw HTML, and every http, https, ftp and www. address comes back one per line, duplicates removed. Here's how the matching behaves, what it cleans up for you, and how to turn the raw list into an audit you can act on.

Where extracting links from documents fits in an audit

Inventory first, judgment second.

Every serious content review begins with an inventory. Before you can ask whether links are broken, outdated, or pointing at competitors, you need to know what links exist. That's easy on a live page and surprisingly annoying in documents: Word drafts, Google Docs exports, markdown files, and newsletter archives all mix addresses into prose where no crawler can see them.

Content that hasn't shipped yet is the highest value case. Catching a dead reference or a staging address in a draft costs nothing; catching it after publication costs a correction. Writers who paste their draft into the extractor before handoff get a checklist of every source in about two seconds, and editors can compare that list against the citations the piece claims to make.

Archives are the other rich vein. A year of newsletter sends, exported as text, condenses into a single deduplicated list of every destination you've ever pointed readers at, which is the raw material for spotting retired campaigns, expired promotions, and partners you no longer work with.

A worked URL extraction with dedupe in action

Three mentions, two results.

Input: The launch post (https://example.com/blog/launch?utm_source=nl) links to https://example.com/docs, and https://example.com/blog/launch?utm_source=nl appears again in the footer. Output: two lines, https://example.com/blog/launch?utm_source=nl and https://example.com/docs. The repeated launch address collapsed into one entry, the query string survived intact, and the summary line reads Found 3 URLs, 2 unique.

Look at what didn't come along: the closing parenthesis after the first link and the comma after the second. The extractor trims sentence punctuation from the end of each link and drops a closing bracket that has no opening partner, so an address like https://en.wikipedia.org/wiki/Mercury_(planet) keeps its brackets while a link wrapped in parentheses loses the stray one. It also turns & in HTML source back into &, so query strings copied from markup stay usable.

Http vs https in extracted lists

Scheme mix is a finding, not noise.

The extractor captures both schemes and treats them as different strings, so http://example.com/page and https://example.com/page show up as two separate lines even though they usually resolve to the same content. A www. address is kept as written, without a scheme, so www.example.com/page is a third line. Don't merge them blindly; the split is telling you something.

A cluster of http entries in a modern document almost always marks old content: links copied from ancient posts, templates that predate a TLS migration, or sources that redirect on every click. Flag every plain http address as an upgrade candidate in your audit sheet. It's the cheapest fix on the whole list and it removes a redirect hop for every reader.

URL grabbing mistakes that skew an audit

Know what the pattern can't see.

The matcher is fast and consistent, but a link audit built on it needs to account for its blind spots.

  • Bare domains are off by default. Addresses starting with www. are always found, but example.com/docs appears only when you tick Also find bare domains, and even then only for common endings like .com, .org, .io and .in, so a file name like report.pdf isn't mistaken for a link.
  • Relative paths don't exist here. A docs page full of /guides/setup style links extracts nothing, since those only become addresses in the context of a site.
  • Near duplicates stay separate. Dedupe compares exact text, so example.com/page and example.com/page/ are two entries. Normalize trailing slashes and case before comparing lists.
  • Dedupe hides frequency. If you need to know that one address appeared nine times, untick Remove duplicates to keep every occurrence in the order it appeared.

Turning the list into audit decisions

From raw lines to decisions.

First, set Sort to Group by domain, or read the Domains list under the options. Grouping by domain splits your audit naturally into internal links, familiar externals, and the long tail of one off sources that need individual review. Second, leave Remove ?query and #fragment unticked so tracking parameters stay visible; inconsistent utm tagging across a campaign is exactly the kind of finding an audit exists to catch.

Third, download the list with the .csv button and timestamp the file. An extracted inventory is a snapshot, and pairing it with a date means the next audit can diff against it instead of starting from zero.

URL Extractor or Link Extractor: pasted text or live pages

Two tools, one distinction.

The names are close, so here's the line between them: the URL Extractor works on any text you paste and never touches the network, while the Link Extractor in the SEO category reads a live page or its HTML and lists each link with its anchor text, internal or external type and rel values. Draft in a doc? Paste it here. Page already on the web? Point the Link Extractor at it.

Around the edges, Find and Replace can add https:// to the www. entries so every row in your sheet is a clickable link, and the Email Extractor runs the same paste for people instead of pages when a document references both.

Try it now

Open URL Extractor

The tool is one click away. No sign up, no upload, no payment.

Open URL Extractor