Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

URL Extractor

Every link in a page’s source or any text — cleaned up, counted and filterable.

Text No upload Works offline Free, no sign-up

Only the text you paste is read — no page is downloaded and nothing is uploaded.

Optional. Turns links such as /about into full addresses and lets you list internal or external links.

Options

Find
Clean up
Needs the page address.

Separate domains with commas or spaces. Subdomains count too: example.com also matches blog.example.com.

Next steps

About the URL Extractor

Paste a web page’s HTML source, an email, a document, a spreadsheet export or any text, and every link in it is listed once with how often it appears. From HTML it reads href, src, srcset, action, lazy-loading data-src attributes, CSS url(…), meta refreshes and links escaped inside scripts and JSON (WordPress writes https:\/\/…); from text it finds http://, https://, ftp:// and www. addresses and, if you want, bare domains such as example.com.

Links are normalised so the same address written differently counts once, relative links such as /about can be completed with the page’s address, and you can filter by domain, by file type (pages, images, PDFs and documents, media, downloads) or to internal or external links, and strip tracking parameters such as utm_source. Nothing is fetched: the tool reads only the text you paste, in your browser.

How to use it

  1. Paste the text or HTML source, or open a file (.html, .txt, .csv, .xml, .md and other text files). To get a page’s source, open it in your browser and press Ctrl+U (⌘⌥U on a Mac).
  2. For HTML, enter the page address so relative links such as /about or images/logo.png become full URLs.
  3. Choose the clean-up — remove #fragments, remove tracking parameters, readable international addresses — and the filters: link type, internal or external, domains.
  4. Pick the order and copy the list, download it as text or CSV, or copy just the unique domains.

Examples

Links from HTML with a page address
Input
<a href="/pricing">Pricing</a> <img src="img/logo.png" srcset="img/[email protected] 2x">
Result
https://www.example.com/pricing
https://www.example.com/img/logo.png
https://www.example.com/img/[email protected]

With https://www.example.com/ entered as the page address.

Links in text, with punctuation and brackets handled
Input
Read https://en.wikipedia.org/wiki/URL_(disambiguation), then www.example.org/news.
Result
https://en.wikipedia.org/wiki/URL_(disambiguation)
https://www.example.org/news

The comma and the full stop end the links; the brackets that belong to the Wikipedia address stay.

Duplicates and tracking parameters removed
Input
https://Shop.Example.com/item?id=7&utm_source=mail
https://shop.example.com/item?id=7#reviews
Result
https://shop.example.com/item?id=7

With “Remove tracking parameters” on (and #fragments removed, the default), both lines are the same page.

Common uses

  • Listing every link, image and script on a page for an SEO or migration audit.
  • Collecting all PDF or image links from a page’s source for download lists.
  • Cleaning a list of campaign links: dedupe them and strip utm_ and click-ID parameters.
  • Finding every external domain a page or a document links to.

How links are normalised

Each link is parsed with the same rules browsers use (the WHATWG URL Standard): the scheme and host become lower case, international hosts are converted to Punycode (bücher.de → xn--bcher-kva.de), default ports such as :443 are dropped, ./ and ../ are resolved and an empty path becomes /. On top of that, percent-encodings are written in capitals and encoded letters and digits are decoded, as RFC 3986 section 6.2.2 recommends, so %7e and ~ count as the same.

  • Remove #fragments (on by default) makes page#a and page#b one page.
  • Remove tracking parameters deletes utm_…, gclid, gbraid, wbraid, fbclid, msclkid, mc_cid, mc_eid, igshid and similar click and campaign IDs, leaving the other parameters exactly as written.
  • Readable addresses shows international hosts and encoded letters (caf%C3%A9) as Unicode (café). Encoded spaces, slashes and invisible characters stay encoded, so the address still works when pasted.

Where a link ends in plain text

Links in text end at a space, at characters that cannot appear in an address (< > " { } |) and at the punctuation of running text: curly quotes and guillemets (“ ” ‘ « »), an ellipsis (…), an em dash (—) and the full-width punctuation of Chinese, Japanese and Korean (。,、:「」). Like GitHub’s autolinks, trailing punctuation — . , : ; ! ? * _ ~ and quotes — is not part of the link, and a closing bracket at the end belongs to the link only if the link also contains its opening bracket, so (see https://example.com/page) and https://en.wikipedia.org/wiki/Foo_(bar) both come out right. Links joined by commas (https://a.com,https://b.com) are split, and a link inside another link — a Wayback Machine address or a ?url= redirect — counts once, as part of the outer one.

Addresses written without http:// are listed with https://. Bare domains (example.com with no www.) are off by default: they must end in a real top-level domain, but file names such as readme.md or setup.py also qualify, because .md (Moldova) and .py (Paraguay) are country domains.

Internal and external links

With a page address entered, a link is internal when its host is the page’s host or a subdomain of it (or the other way round), ignoring www.: for www.example.com, links to example.com and blog.example.com are internal, and links to example.org are external. Relative links are always internal.

Limitations

  • Only the text you paste is read. Links that a page adds with JavaScript after it loads are not in its source; to include them, copy the live HTML from your browser’s developer tools (in Chrome or Edge: Elements panel, right-click the <html> element, Copy, Copy outerHTML) and paste that.
  • Relative links found in JavaScript code, and links inside images or PDFs, are not found. For a PDF, get its text first with PDF to Text.
  • Two different URLs can show the same page (for example with and without a trailing slash). They are listed separately, because servers may treat them differently.
  • Where a link’s path runs straight into Chinese, Japanese or Korean words with no space or punctuation in between (…/aboutをご覧ください), those words are kept as part of the path, because paths in those scripts are common (for example Wikipedia titles). A host that runs into words is cut correctly.
  • Readable addresses show international hosts in their own letters, so a look-alike such as аррӏе.com, written with Cyrillic letters, reads like apple.com. Keep that option off when checking links for phishing: the Punycode form (xn--80ak6aa92e.com) shows the difference.
  • Files up to 50 MB can be opened; very large pages may take a few seconds.

Privacy

Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.

Frequently asked questions

How do I extract all links from a web page?

Open the page, view its source (Ctrl+U, or ⌘⌥U on a Mac), select all and copy it, then paste it here and enter the page’s address so relative links become full URLs. The tool never downloads the page itself.

How do I get only the image or PDF links?

Choose Images or Documents under Link type, or Custom extensions and type, for example, pdf, docx. The type is taken from the file extension at the end of the path.

Can it remove utm parameters from my links?

Yes. Tick Remove tracking parameters: utm_source, utm_medium, gclid, fbclid and similar parameters are deleted and the remaining ones are kept exactly as they were. Links that differed only in tracking parameters then count as one.

Why do some links show “relative”?

Links such as /about or images/logo.png only make sense relative to the page they came from. Enter the page’s address and they are completed. If the HTML contains a <base href> tag, it is used automatically.

Does it check whether the links work?

No — it only reads text and never connects to the sites. To see where a link leads, use the Redirect Checker.

Is my text uploaded?

No. Everything happens in your browser; nothing you paste or open is sent to a server.

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.