URL Extractor

Extract URLs from text or HTML and get a clean list of links. Trailing punctuation is trimmed, duplicates are removed, and you can keep only the links of one domain.

Updated
Runs in your browser. Your data is not uploaded.

How to use the URL Extractor

  1. Paste text or HTML into Text or HTML, upload a file, or press From URL to load a page.
  2. Set Read the input as (automatic detection works for most input) and tick the options you need, such as Domains only or Find bare domains. To keep one site's links, fill in Only links on this domain.
  3. Copy the list under Links found or download it as a .txt file.

How it works

  • Plain text: the extractor finds anything that starts with http://, https://, or www. and runs to the next space or quote.
  • Clean-up: a period, comma, or other sentence punctuation at the end is removed, and so is a closing bracket that has no opening bracket inside the link. (see https://example.com/page). gives https://example.com/page, but https://en.wikipedia.org/wiki/Mercury_(planet) keeps its bracket.
  • HTML: the markup is read with your browser's HTML parser, never displayed, and the href, src, and srcset values that are absolute links are collected, along with any links in the text.
  • Bare domains: when ticked, words like example.com count too, but only if they end in a known top-level domain. The list is loaded only when you turn this on. Email addresses are skipped.

Examples

  • Docs: https://example.com/docs/start (see the FAQ). → https://example.com/docs/start.
  • <a href="https://shop.example.com/cart">Cart</a> → https://shop.example.com/cart.
  • With Find bare domains: status page is status.example.com → status.example.com, while readme.md is ignored as a file name.
  • With Domains only, https://www.example.org/blog and www.example.org/about both become example.org.

Limitations

  • Relative links in HTML, such as /about, are resolved against the page address when you use From URL. In pasted HTML they are skipped, because the page they belong to isn't known.
  • From URL works only when the site allows browsers to read it (CORS), and reads up to 5 MB. Many sites refuse; then open the page, view its source, and paste it.
  • Bare domain detection uses a list of country-code and common generic TLDs, not the full Public Suffix List. Rare new TLDs may be missed, and a file name like setup.io may be read as a domain.
  • Links are found as written. Shortened links are not expanded, and nothing is visited.

Frequently asked questions

How do I extract all links from a web page?

Open the page, view its source (Ctrl+U or Cmd+Option+U), copy it, and paste it here. The extractor reads the href and src attributes and lists every absolute link.

How do I get only the domains?

Tick Domains only. Each link is reduced to its host name without www., and with Remove duplicates on you get one line per domain.

Are the links visited or checked?

No. Extraction happens in your browser, and none of the links are opened, so it is safe to paste suspicious emails.

Often used together with the URL Extractor.