SOURCE → STRUCTURE → HANDOFF
Free URL, Link and Domain Extractor
Turn supplied text, HTML, Markdown or structured source into a clear report of URLs, link occurrences, hostnames and registrable domains. Choose automatic detection or an explicit format, review source evidence, and filter the results before copying or downloading the view you need.
Analysis runs locally in the isolated tool workspace. It extracts information from the content you provide without visiting the extracted URLs.
Open the URL, Link and Domain ExtractorUse embedded workspace
Paste your source, choose its format and select Extract. You can also choose supported local UTF-8 files. Review each source’s mode and generic label; a batch uses one shared processing budget rather than a separate allowance for each file.
Default input limit: 512 KiB per run, shared across supplied sources. A run processes up to 500 candidates and may return a partial report when a processing limit is reached. Check the active limits in the workspace. Copy results or download CSV, JSON or TXT.
YOUR SOURCE. YOUR NEXT STEP.
Extract and review your source
The embedded workspace analyzes what you supply. It does not visit extracted destinations.
The workspace below processes supplied source locally. Open the isolated extractor in a separate tab if you need more room.
Need more room? Open the isolated workspace.
Choose the right input mode
The input mode determines what the extractor looks for. Select the format that matches your source rather than expecting every type of reference to be treated as a website link.
| Input mode | What to supply | What it extracts |
|---|---|---|
| Text | Article text, copied notes or other plain text | Explicit http://, https:// and www. URLs. A www. address uses the chosen default scheme—HTTPS initially—and is flagged as an assumption. |
| HTML | Supplied HTML source or a copied HTML fragment | HTTP(S) destinations from a[href] and area[href], with raw href, label provenance, rel, target and title. Relative references need an explicit base URL. Optional resources and non-HTTP references have separate views. |
| List | One item per non-empty line | HTTP(S) URLs or bare hostnames such as example.org. A bare hostname uses the chosen default scheme—HTTPS initially; example.org/path needs an explicit scheme. |
| Auto | Supported source whose format you want detected | Shows the detected mode; switch to an explicit mode when the source is ambiguous. |
| Markdown | Supplied Markdown links and image references | Inline, reference-style and autolink syntax using the supported parser; fenced code is not link evidence. Image references are resources, not navigation links. |
| CSV | Comma-separated UTF-8 records | Parses quoted fields and multiline cells. Choose all columns or a column index, and set whether the first record is a header. |
| JSON | A valid JSON source document | Decodes and scans string values, including escaped URLs. Values remain bounded; object keys are not a website lookup. |
| XML | Supplied XML without a DTD or entity declaration | Parses supported text and attribute values. External entities are rejected; no resource is fetched. |
| Mixed | Text containing explicit URLs, bare hosts or email addresses | A separate deliberate domain-discovery mode. Email-derived hosts are domain evidence, not a fetched URL or an email verification. |
In Text mode, an email address or an unprefixed domain such as example.org is not a URL match. Use List mode for one bare hostname per line, or choose Mixed mode deliberately to review supported bare-host and email-derived domain evidence.
Local files must contain supported UTF-8 text, HTML, Markdown, CSV, JSON or XML. CSV text is not an Excel workbook decoder. The tool does not decode PDFs, Word documents, Excel files, email MIME attachments or ZIP archives.
Choose a view for the question you are answering
Unique URLs: which distinct addresses occur?
The URL view lists normalized HTTP(S) addresses and their occurrence counts. Identical normalized URLs are combined, so you can separate a repeated address from several different destinations.
Normalization uses standard URL parsing. The hostname and scheme are normalized, default ports and dot segments are handled, and path case is preserved. HTTP and HTTPS, query variations, fragments and trailing-slash differences can remain separate addresses.
HTML links: where is each link represented?
The link-occurrence view retains individual references. Review the raw href beside the resolved destination, visible text and available label fallback, rel, target, title, tag and source position. A supplied-HTML structural path identifies parsed document position, not an exact original byte offset. Long retained values have explicit bounds and truncation information.
This view is useful when one URL appears several times with different labels or rel attributes. An image-only anchor can use its nested image alt text as a fallback; supported aria-label and title fallbacks are identified separately from visible text. An empty-label flag is evidence for review, not an automatic accessibility or ranking verdict.
Hostnames: which hosts appear?
The hostname view groups addresses by complete hostname. For example, www.example.co.uk and docs.example.co.uk are different hosts. Counts help you see how often a host appears in the retained source.
Registrable domains: which suffix-based groups appear?
The registrable-domain view groups hosts using Public Suffix List rules, including private suffixes. This avoids treating every final two-label combination as a domain root.
For example, www.example.co.uk groups under example.co.uk. Hosts such as team.github.io and other.github.io can remain separate registrable-domain groups because the private suffix rules matter.
These groups describe domain structure. They do not identify a site’s owner or establish that it is registered, reachable or suitable for a campaign.
Resources and metadata: what else is referenced?
Enable optional resource extraction to review supported images, scripts, stylesheets, media and form references separately from navigation links. Canonical, hreflang and other supported link-element references retain their supplied role. A reference is not proof that the resource loads or that a search engine accepts its metadata.
Non-HTTP references: what was not treated as a webpage?
Supported special schemes are classified in a separate view. A mailto or tel reference is not a website URL, and a script-like value is never executed. These references do not silently inflate HTTP(S) URL totals.
How to extract URLs, links and domains
- Supply your source. Paste a supported format or choose a local UTF-8 source file. Keep meaningful source content and inspect any format error before extracting.
- Select the matching mode. Use Auto detection as a starting point, then confirm its choice. HTML, Text and List keep their own rules; Markdown, CSV, JSON, XML and Mixed modes have the boundaries described above.
- Set a base when needed. For relative HTML references, enter the actual HTTP(S) page URL that provides the intended context.
- Review the result and diagnostics. Check assumptions, rejected candidates, counts, unresolved references and whether the report is complete or partial.
- Copy or download the appropriate scope. Keep the full-report JSON or full-view CSV when you need every retained result. For a smaller handoff, choose the filtered or selected-row scope and review its row count before copying or downloading.
Leave query and fragment removal off when those parts may matter. Review the effect of an optional cleanup before using the resulting list elsewhere.
Example: repeated URLs in supplied text
Consider this input in Text mode:
Visit https://example.com/Guide?ref=one#intro, and www.example.org/help.
Again: https://example.com/Guide?ref=one#intro.
A bare example.net and email [email protected] are not text-mode links.With the default HTTPS assumption and cleanup disabled, the extraction engine returns two unique URLs from three accepted occurrences:
| Normalized URL | Occurrences | Interpretation |
|---|---|---|
https://example.com/Guide?ref=one#intro | 2 | The repeated address is combined; the query, fragment and path case remain. |
https://www.example.org/help | 1 | HTTPS is assumed for the supplied www. address and recorded in the report. |
The sentence punctuation is excluded from the matched addresses. The bare domain and email are not extracted in this mode. This example was checked against the current downloaded extraction engine; it does not involve visiting either destination.
Resolve relative HTML links with the correct base
An HTML link such as href="/about/" does not contain its own hostname. Supply an explicit base to resolve it meaningfully.
With the base https://example.com/guides/draft/, standard URL resolution produces:
| Supplied href | Resolved destination | Relationship to the base hostname |
|---|---|---|
/about/ | https://example.com/about/ | Internal |
../resources/ | https://example.com/guides/resources/ | Internal |
#sources | https://example.com/guides/draft/#sources | Internal |
https://docs.example.com/start | https://docs.example.com/start | External |
The default relationship check compares exact hostnames, so a different subdomain is external in this example. Optional www-alias or registrable-domain policies change that classification deliberately; read the selected policy in the report. Without a supplied base, relationships are unknown and relative references are unresolved. Embedded HTML base tags are ignored rather than silently overriding your chosen context.
This table illustrates the resolution and classification rules. It is not a response-status check.
Keep useful URL differences when cleaning a list
Duplicate removal and URL cleanup answer different questions. By default, the extractor preserves differences that may identify distinct resources.
For example, these remain three different normalized URLs:
https://example.com/page?x=1#one
https://example.com/page?x=2#two
https://example.com/pageChoosing both remove query strings and remove fragments combines them into one URL with three occurrences. The report records the removals. That can be useful for a broad page-path review, but the original parameters or fragments may be important for products, attribution or document navigation.
For selective cleanup, choose the tracking-parameter keys you want removed instead of deleting the whole query. Inspect the original-to-output values and recorded changes; preserve functional keys such as product identifiers. Optional scheme or trailing-slash formatting is a rewrite, not proof of working HTTPS or the website’s canonical URL.
Inspect URL components without hiding the original
Review the parsed scheme, port, path, query and fragment alongside the normalized address. Query entries keep repeated keys and empty values. These fields explain the supplied URL’s structure; they are not an HTTP response or a canonical check.
Understand hostname, suffix and domain results
The Public Suffix List contains rules for suffix boundaries. It allows the tool to handle structures such as .co.uk and private hosting suffixes more accurately than taking the last two labels of every hostname.
Private suffix rules are enabled by default. You can choose an ICANN-only grouping policy when that matches your task; hosted-tenant groups may then combine differently. Optional www removal is an output-formatting choice, separate from relationship classification. The report records the policy and bundled library version. Its suffix data is a pinned snapshot, not a live lookup.
IP addresses, localhost and unknown-suffix hosts may have no registrable-domain result. Internationalized hostnames retain normalized ASCII/punycode values, with Unicode display fields alongside them for review. Choose full host, registrable domain, subdomain, public suffix or final-label output; the final label is not the same as a registrable domain. A missing root or a recognized suffix is a classification result, not a DNS lookup, registration check or availability decision.
The Public Suffix List’s explanation describes this distinction and why the list should not be used as proof that a domain exists.
Inspect anchor text and rel values in context
Use the HTML-link view to review the actual labels and attributes supplied in your source. Check whether a label explains its destination, whether a URL appears with several different labels, and whether a supplied rel value includes tokens such as nofollow, sponsored or ugc.
These attributes describe how the link is qualified. They do not provide a guaranteed measure of ranking value, and an absent nofollow token does not promise crawling or ranking benefit. Google’s guidance on qualifying outbound links explains the intended use of these values.
For more context around the wording of a link, use the Anchor Text Optimizer. Review a public page’s outbound-link counts and supplied rel information with the Outbound Link Counter.
Filter, sort and export an intentional subset
Focus on internal or external relationships, supplied rel tokens, anchor wording or empty labels. Hostname filters distinguish exact matches from a deliberate domain-suffix match; a substring in an unrelated hostname is not enough. Public-suffix filters use the report’s suffix field, not a guess from the final two labels.
Sort by source order, address, hostname or frequency with stable ordering for ties. Counts distinguish retained occurrences, unique URLs and grouped hosts or domains. Changing a display filter does not make a partial scan complete.
Use the full export actions for every retained row, or choose an explicit filtered or selected-row scope for a smaller list. Review the scope, row count and columns first. Copy has a selectable-text fallback when clipboard access is denied. TXT, CSV and JSON serve common handoffs; supported TSV, Markdown, HTML, bookmark and JavaScript-array outputs apply their own escaping. A generated HTML file does not run inside the tool, and no native Excel workbook is claimed.
Use the extracted list in a wider editorial or SEO workflow
Review an article before publication
Extract links from your article’s supplied HTML, check repeated destinations and review labels. Prepare or edit supported source formatting in the HTML and Markdown Editor, then rerun extraction on the revised source.
HTML source and rendered HTML can differ. A copied source document may not include links inserted later by JavaScript. The extractor analyzes what you supply and does not run the site’s application to discover additional links.
Check live destinations separately
An extracted URL is not proof of a working destination. Use the Broken Link Finder for its supported public-page scan and response diagnostics. Check the source page and affected URL there; the extractor’s download is a manual handoff, not an automatic connection between the tools.
Review internal-link opportunities across a site
One pasted page cannot establish which pages lack internal links across an entire website. Use the Internal Link Opportunity Finder for its bounded crawl and supported site-level comparisons. Treat incomplete discovery and orphan candidates as items for review.
Prepare a publisher shortlist
Turn a supplied publisher URL list into hostname and registrable-domain groups to reduce repeated entries. Keep private-hosting groups separate where the suffix rules require it, and review each site rather than assuming a root-domain count equals a publisher count.
Use the Bulk Publishers Availability Checker to check its supported domain inputs against EduGuestPost’s configured private catalog. A catalog match is separate from domain registration, publisher quality or a guaranteed placement.
Review the editorial standards before pursuing a placement. If you need help with publisher research, pitches and coordination, explore guest posting and blogger outreach.
Local analysis and export privacy
Your supplied text or file content is analyzed inside the isolated extractor workspace. That workspace is configured to block network connections during analysis, and supplied HTML is parsed without inserting it as live page content. Extracted values appear as text rather than links that automatically open.
The surrounding EduGuestPost website still makes normal website requests. The local-processing description applies to the tool’s handling of your supplied source, not to every request made by the parent page. See the site’s privacy policy for broader website information.
Reports do not include your original source document or automatically include local filenames. Generic source labels identify batch inputs unless you deliberately choose a label. Extracted URLs, raw references, labels, attributes and source metadata can still contain private information. A scoped export is not an anonymized export; review it before sharing.
Know the limits of each report
The default workspace limits are 512 KiB of UTF-8 input and 500 processed candidates per run, shared by up to five supplied sources. Candidates include duplicates and rejected entries, so this is not a promise of 500 accepted URLs. Additional URL-length, HTML-node, time and output limits can also end a scan early. Check the active configured limits in the workspace.
A partial report contains the retained results and the reason processing stopped. It does not establish that the remaining source has no more URLs. Split a larger source into manageable parts if needed, and account for duplicates when reviewing several reports.
Results display in pages of 25 rows. The existing full-view CSV action exports every retained row in its view; full-report JSON retains the complete report. Separate scoped actions can export filtered matches or selected rows, with an explicit row count. Pagination alone does not limit an export. The default downloadable output limit is 2 MiB; the active workspace limit is shown.
The extractor does not fetch websites, crawl a site, test HTTP status, resolve redirects, check domain ownership, identify backlinks pointing to your site or confirm Google indexing. Its counts describe the supplied source within the report’s stated scope.
Frequently asked questions
Is the URL, Link and Domain Extractor free?
Yes. The current local workspace can be used without an API key or paid extraction subscription, within its displayed limits.
Can I extract links from a website URL?
The current tool analyzes supplied content. Copy supported text or HTML from the page and choose the appropriate mode. Entering a URL does not instruct it to retrieve that website.
Can I extract bare domains from text?
Use List mode with one bare hostname per line, or Mixed mode for supported domain and email-derived-host discovery. Text mode still looks for explicit HTTP(S) and www. URLs and does not silently change its rules.
Does it keep subdomains?
Yes. The hostname view preserves complete hostnames, while the registrable-domain view groups them using the supported suffix rules. Private suffixes can keep different hosted tenants in separate groups.
Does it identify duplicate URLs?
It combines identical normalized URLs and records occurrences. It does not establish that differently written URLs lead to the same live page.
How are relative HTML links handled?
Supply an explicit HTTP(S) base URL. Without one, relative references cannot be resolved and are reported. Embedded base tags are ignored.
Can it show internal and external links?
Yes. With a supplied base, the default policy compares exact hostnames, so a different subdomain is external. Optional www-alias or registrable-domain policies are deliberate alternatives recorded in the report. Without a base, the relationship is unknown.
Can it identify nofollow links?
The HTML-link view reports supplied rel values, so you can inspect nofollow, sponsored or ugc tokens. It does not determine how a search engine will treat an individual link.
Can I upload PDF, Word or Excel files?
No native PDF, Word or Excel decoding is included. Supported UTF-8 text, HTML, Markdown, CSV, JSON and XML can be loaded locally. CSV is a text-record format, not an Excel workbook. Convert other document formats into supported source first.
Can I copy results or download Excel files?
You can copy a supported list or download TXT, CSV and JSON, with additional scoped formats shown in the workspace. Clipboard denial has a selectable-text fallback. Native Excel workbooks are not included; CSV can be imported into a spreadsheet as text.
Does CSV include only the rows I filtered on screen?
Choose the scope deliberately. Full-view CSV keeps all retained rows in that view; the separate scoped export can use filtered matches or selected rows. Review its row count before saving. The 25-row page only controls display and does not silently limit a download.
Does a complete result prove all links are valid or indexed?
No. Complete means the supported input scan finished within its processing boundaries. Reachability, registration and Google indexing require different checks.
Turn your source into a usable link report
Choose the right mode, supply a base for relative HTML links, and inspect the report’s counts and diagnostics before downloading your results.
