My Tool Studio
SEO Tools·4 min read

Running an Internal Link Audit With a Link Extractor

Running an internal link audit is the highest-yield SEO work most sites never do, because the data gathering feels like a chore. It doesn't need to be. The Link Extractor pulls every anchor from a URL in one request, splits internal from external, shows each rel value and counts nofollow links, which turns the chore into a paste-and-export loop. This guide covers the whole audit: what the split counts mean, how nofollow works in practice, how to find orphan pages with nothing but exports, and the mistakes that make link audits lie to you.

Link Extractormytoolstudio.com › tools<title>…</title><meta name="description">SEO ready

The strong page your internal links ignore

Authority flows where links point.

Every site has one: a page that converts well, targets a valuable query, and sits three clicks deep with two internal links pointing at it, both from footers. Meanwhile, the blog post that went mildly viral a few years ago still collects forty inbound internal links it no longer deserves. Authority flows where links point, not where you wish it went.

You can't fix a distribution you've never measured. Measuring starts with extraction: one page at a time, every anchor, resolved and categorized.

Running an internal link audit, page by page

Pick the ten URLs that matter most to revenue, plus your homepage. For each one, paste the URL, press Extract links, and read the badges: total, unique, internal, external, nofollow. Then press CSV and name the file after the source page. Each row holds the URL, anchor text, type, rel value, nofollow flag, position on the page, and how many times the link repeats.

Ten minutes later you have an adjacency map of your site's core. The questions nearly answer themselves. Which money pages appear in nobody's list? Which pages appear in every list because they sit in the nav, telling you nothing? Where does a high-authority page link out to only five body targets, leaving room for three more?

A worked link extraction with real counts

Input: one blog URL.

Extract a typical blog post and the badges might read: Input: one URL. Output: 148 total, 96 internal, 52 external, 7 nofollow. Those numbers look rich until you pick the Internal pill and recognize the shape: roughly sixty of the internal links are navigation, footer, and category boilerplate that appear on every page of the site. The position menu makes this concrete; switch it to Main content and only the links inside the article remain.

The body itself links to five pages. That's the real number for audit purposes, and it's why the anchor text shown under each URL matters. Rows reading 'Learn more' and 'Click here' are links spending their descriptive power on nothing, while 'cold brew ratio guide' tells users and crawlers exactly what's on the other end.

Nofollow, sponsored and ugc in the extracted list

Nofollow used to be a bright line: it passed nothing. Since 2019 Google treats rel=nofollow as a hint rather than a directive, alongside rel=sponsored and rel=ugc. Each row shows its full rel value as a small badge, and the Nofollow, Sponsored and UGC pills show each group on its own. For audits, the nofollow view matters most on outreach: when a mention you earned turns out to be nofollowed, the value calculation changes even if some signal may still pass.

On your own site, internal nofollow is almost always debris from old PageRank sculpting advice. If Internal plus Nofollow shows any entries, investigate, because you're effectively asking Google to distrust your own pages.

Orphan page detection using nothing but exports

Finding orphans is set subtraction. One set is every page you want indexed, which your sitemap already lists. The other is the union of internal links you've extracted. Pages in the first set that never appear in the second are orphans: reachable through the sitemap, invisible to users, starved of authority.

The page-by-page version catches the worst cases fast. Extract your highest-traffic hub pages, and any money page absent from every one of those lists is functionally orphaned, whatever a full crawl might eventually add.

Link audit mistakes that skew the picture

Watch for five distortions:

  • Treating navigation and footer links as endorsements, which makes every page look equally linked.
  • Auditing only the homepage, whose link profile resembles no other page on the site.
  • Ignoring anchor text, so twenty 'Learn more' links pass as healthy internal linking.
  • Counting a redirected target as a working link when the old URL burns a hop on every click and crawl.
  • Reading external link counts as leakage, when citing good sources is normal and expected.

Tips for keeping link exports usable

Filter before you copy; Internal plus Remove duplicates gives Copy URLs a clean set with no manual work. Label every export with its source URL straight away, because fifty unlabeled lists are indistinguishable by lunchtime.

Also note the display step. The list shows 200 links at a time, but Copy URLs, CSV and Excel include every link that matches the filters, so trust the export over your scrolling.

For pages that aren't live yet, or for an email template, switch to Paste HTML and drop in the source. Add the page URL in the optional field so relative links turn into full addresses and internal links are recognized as internal.

Link Extractor, Broken Link Checker, or both

The Link Extractor inventories what a page links to, and its Check status codes button tests up to 300 of those targets for 404s and redirects. The Broken Link Checker does the same across a whole site, following internal links from page to page. A serious audit uses them in sequence: extract to map one page's structure, then crawl to verify every destination. And when the audit widens from anchors to media, the Image Extractor from URL applies the same paste-a-URL pattern to every image on the page.

Try it now

Open Link Extractor

The tool is one click away. No sign up, no upload, no payment.

Open Link Extractor