Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

Link Extractor

Every link on a page, with its anchor text and rel values — ready to filter and export.

Network Uses live data Free, no sign-up

Page to read

MySmartCoPilot’s server fetches the page (up to 2 MB, following redirects) and sends back only its links — never the page itself.

Also list

Next steps

About the Link Extractor

See every link on a web page in one table: the resolved URL, its anchor text, the rel values search engines care about (nofollow, sponsored, ugc), whether it opens in a new tab, how often it appears and whether it stays on the site. Filter internal, external and subdomain links, search them, and export the list as CSV.

Enter a URL and MySmartCoPilot’s server fetches the page and extracts the links with Cloudflare’s HTMLRewriter — the page itself is never sent back. Or paste the HTML source and the links are read entirely in your browser. Relative addresses are resolved the way browsers do it, including a <base href> in the page.

How to use it

  1. Choose From a URL and enter a page address, or Paste HTML and paste the page source (optionally with the page’s URL, so relative links can be resolved).
  2. Tick Also list to include images, scripts or <link> tags as well as links, iframes and image-map areas.
  3. Press Extract links, then filter by Internal, External, Subdomains or Other (e-mail, phone, JavaScript), search, or show only nofollow/sponsored/ugc links or duplicates.
  4. Press Copy URLs for a plain list of what is shown, or download a CSV of the unique links or of every occurrence.

Examples

Duplicates, rel values and a same-page link
Input
<a href="/about">About us</a>
<a href="/about" rel="nofollow">About</a>
<a href="https://partner.example.org/" rel="sponsored" target="_blank">Partner</a>
<a href="#top">Top</a>
(page URL https://www.example.com/blog/)
Result
4 links in the page · 3 unique URLs · 1 internal · 1 external · 1 same page
https://www.example.com/about — “About us”, “About” · 2 times · nofollow 1/2
https://partner.example.org/ — “Partner” · sponsored, _blank
Pasted HTML with a base element
Input
<base href="/en/">
<a href="pricing" rel="nofollow">Prices</a>
Result
https://www.example.com/en/pricing — “Prices” · nofollow (page URL https://www.example.com/blog/)

With a <base href>, relative links resolve against it, not against the page address.

Common uses

  • Auditing internal linking: which pages a page links to, with what anchor text, and how often.
  • Checking that paid and affiliate links carry rel="sponsored" and user-generated ones rel="ugc".
  • Collecting the outbound links of a page for a broken-link or link-building check.
  • Pulling the links out of an HTML e-mail or a saved page without opening it in a browser.

What is extracted

  • <a href> links with their anchor text — including the alt text of an image inside the link, which Google uses as anchor text — and <area href> links of image maps with their alt text.
  • <iframe src> embeds.
  • Optionally <img src> and every srcset candidate, <script src>, and <link href> with its rel (stylesheet, canonical, alternate with hreflang, icon…).

Links inside HTML comments, scripts and iframe fallback content are not links and are left out. When a link has no text, its aria-label or title is used, and links without any are counted — they give screen-reader users and search engines nothing to go on.

How addresses are resolved and grouped

Relative addresses are resolved as browsers do (RFC 3986 §5 through the WHATWG URL parser). The first <base> element with an href sets the base for every link in the page; a base that cannot be parsed, or that uses data: or javascript:, is ignored (HTML Standard, the base element).

Internal means the same host as the page (with or without www.), Subdomains another host of the same registered domain by the Public Suffix List — so two sites on github.io count as external to each other — and Same page a link to a fragment (#…) of the page itself.

nofollow, sponsored and ugc

Google asks sites to mark paid links and adverts with rel="sponsored", links in comments and forum posts with rel="ugc", and other links it should not associate with the site with rel="nofollow"; values can be combined, separated by spaces or commas (Google Search Central: Qualify your outbound links). Google says such links are generally not followed, but the pages may still be crawled when found another way. The table counts each value per occurrence, so “nofollow 1/3” means one of three links to that URL has it.

Limitations

  • Links added by JavaScript after the page loads are not seen: the server reads the HTML as delivered, without running scripts.
  • Pages larger than 2 MB are read up to 2 MB, and at most 5,000 links are listed.
  • Sites can answer MySmartCoPilot’s server (Cloudflare, as MySmartCoPilotBot) differently from a browser, or block it; paste the page source in that case.
  • Only public http:// and https:// pages on the standard ports can be fetched; private and internal addresses never are.
  • Text is read as UTF-8; anchor text of pages in other encodings may show wrong characters (the URLs are not affected).

Privacy

In “From a URL” mode the address goes to MySmartCoPilot’s server, which fetches the page from Cloudflare’s network and returns only the links; nothing is stored, and the log keeps only the host name, status and time. Pasted HTML is processed in your browser and never uploaded.

Frequently asked questions

Why are some links missing that I can see in my browser?

Menus and widgets built with JavaScript add their links after the page loads, and the extractor reads the HTML as the server sends it. Paste the HTML from your browser’s developer tools (Elements panel → copy the html element) to include them.

What counts as an internal link?

A link to the same host as the page, ignoring a leading www. Links to other hosts of the same registered domain (shop.example.com from www.example.com) are listed as Subdomains, and everything else as External.

Do I still need rel="noopener" with target="_blank"?

Current browsers treat target="_blank" as if rel="noopener" were set, so the opened page cannot control yours. The table still shows target and rel values so you can check them.

Can I export every occurrence instead of unique URLs?

Yes. CSV lists each URL once with its count and anchor texts; CSV, every occurrence lists every link as it appears in the page, with the address as written, its tag, rel and target (and the line number for pasted HTML).

Is it safe to paste HTML from an untrusted page?

Yes. The HTML is read as text by a parser that never runs scripts, loads images or opens links, and it stays in your browser.

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.