← ALL FIELD NOTES

How to Extract Linked Files from Websites Reliably

A practical two-stage ingestion pattern for discovering attachment links, preserving context, validating downloads, and routing files safely.

Most web scrapers begin by extracting the readable text of a page. That is useful for articles, product pages, and documentation—but it misses a common source of primary material: files linked from the page.

A meeting page may contain agendas and minutes. A policy page may link to a PDF, spreadsheet, or archived record. A data portal may put the most useful dataset behind a button labeled “Download.” Treating these as ordinary page text loses both the file and the relationship that explains why it matters.

To extract linked files from websites reliably, use a two-stage ingestion design:

  1. Discover and contextualize candidate links on the source page.
  2. Resolve, validate, retrieve, deduplicate, and route each accepted resource.

The separation is important. A page parser answers, “Which links appear relevant here?” A download worker answers, “What did this URL actually return, and can we safely process it?”

Stage 1: discover links without losing their meaning

Start with hyperlink records, not just rendered text. The HTML anchor element can link to many URL-addressable targets, including pages, files, email addresses, and fragments. Its href is the destination, while its download attribute is a useful signal that the author intends a download. See the MDN anchor element reference.

Do not turn every anchor into a download job. Navigation, social links, mail links, fragment links, thumbnails, and tracking destinations will otherwise swamp the queue. Instead, collect candidates with enough context to make a policy decision later.

A useful candidate record includes:

  • sourcePageUrl: the canonical page where the link was found
  • rawHref: the literal destination from the page
  • absoluteUrl: the resolved fetch URL
  • linkText: normalized visible label, if present
  • downloadHint: the page-provided download filename hint, if present
  • rel: link relationship metadata, if present
  • sectionHeading: the nearest preceding heading or attachment-list label
  • surroundingText: a short, bounded text window
  • domOrder: the link’s position on the page

This context makes a substantial difference. A link labeled “Annual report (PDF)” under a heading named “Downloads” is a different candidate from the same URL found in a site footer. Preserve both records until your policy decides otherwise.

Resolve URLs with a URL parser

Attachment lists often use relative paths such as files/report, root-relative paths such as /documents/report, or parent-directory references. These should be resolved using a standards-based URL parser rather than string concatenation. The URL() constructor resolves a reference against a base URL correctly.

type ParsedAnchor = {
  href: string;
  text: string;
  download?: string;
  rel?: string;
  sectionHeading?: string;
  surroundingText?: string;
  position: number;
};

type FileCandidate = {
  sourcePageUrl: string;
  rawHref: string;
  absoluteUrl: string;
  linkText: string;
  downloadHint?: string;
  rel?: string;
  sectionHeading?: string;
  surroundingText?: string;
  domOrder: number;
};

function discoverCandidates(
  pageUrl: string,
  anchors: ParsedAnchor[],
): FileCandidate[] {
  return anchors.flatMap((anchor) => {
    let resolved: URL;

    try {
      resolved = new URL(anchor.href, pageUrl);
    } catch {
      return [];
    }

    if (!['http:', 'https:'].includes(resolved.protocol)) return [];
    if (resolved.hash && resolved.pathname === new URL(pageUrl).pathname) return [];

    return [{
      sourcePageUrl: pageUrl,
      rawHref: anchor.href,
      absoluteUrl: resolved.href,
      linkText: anchor.text.trim(),
      downloadHint: anchor.download,
      rel: anchor.rel,
      sectionHeading: anchor.sectionHeading,
      surroundingText: anchor.surroundingText,
      domOrder: anchor.position,
    }];
  });
}

This code deliberately does not infer that a candidate is a PDF or an attachment from its URL. URL suffixes are useful hints, but they are neither authoritative nor always present.

Rank candidates before fetching

Use explainable rules before download. For example, raise a candidate’s priority when:

  • the page supplies a download hint;
  • the text or nearby heading contains terms such as “attachment,” “download,” “agenda,” “report,” or “data” for your domain;
  • the URL path has an allowed extension;
  • the link occurs inside a known attachment-list component.

Lower its priority when it is in a header or footer, has no meaningful label, points to an obvious navigation destination, or has a disallowed scheme. Keep the reason codes. They make false positives diagnosable and let you improve ranking without reinterpreting raw pages.

Build the ingestion layer around source context, not filenames alone. If you are evaluating PagePith for page retrieval workflows, create an account and test it against representative source pages.

Stage 2: verify the response, then retrieve it

A candidate URL is only a proposal. The HTTP response determines what you received.

The key response signals are:

  • Final URL and redirect chain: record where the request ended, not only the original link.
  • Status code: distinguish successful retrieval, access denial, and missing resources.
  • Content-Type: identifies the media type of the returned representation.
  • Content-Length: when present, provides an expected byte length.
  • Content-Disposition: may identify an attachment and provide filename or the internationalized filename* parameter.

HTTP defines Content-Type and Content-Length as representation metadata, and a HEAD request can expose metadata without a response body. In practice, treat HEAD as an optional optimization, not a required gate: some servers omit metadata or handle HEAD differently. The relevant HTTP semantics are specified in RFC 9110.

A pragmatic decision flow is:

  1. Attempt a metadata request when your target sites support it.
  2. Apply size, type, destination, and redirect policies.
  3. Fall back to a bounded GET when metadata is unavailable or unreliable.
  4. Stream accepted responses to temporary storage while enforcing a hard byte limit.
  5. Inspect the completed response metadata and calculate a checksum.
  6. Promote the file only after validation succeeds.

This keeps a page crawl from becoming an unbounded download service.

Select filenames from multiple signals

Do not trust the URL path as the filename. A route named download?id=17 may return a spreadsheet, while a .pdf path may return an HTML login page.

Use a precedence policy such as:

  1. a valid server-provided filename* value;
  2. a valid server-provided filename value;
  3. the final URL path basename;
  4. a generated stable name based on a digest and detected media type.

Content-Disposition supports both filename and filename*; RFC 6266 recommends preferring filename* when both are present. The same specification cautions recipients not to let a supplied filename write outside an authorized location and recommends stripping path information. Read the details in RFC 6266.

In implementation terms, sanitize every server-supplied name: remove path separators and control characters, reject empty or reserved names, enforce a length limit, and write only beneath a generated storage directory. The filename is presentation metadata, not a filesystem instruction.

Route files using evidence, not extensions

Once a file is retrieved, route it using the combined evidence from the response and the bytes you actually stored:

EvidenceAppropriate use
Content-TypeInitial processor selection
Content-DispositionFilename and attachment intent
Final URLProvenance and fallback naming
File signature or parser probeValidation before expensive processing
Source-page contextRelevance, labels, and downstream attribution

For example, an endpoint may claim application/pdf but return a small HTML access page. A lightweight file signature check or parser probe can catch that mismatch before a PDF extraction task is queued. Retain the original media type claim and your detected result separately; disagreement is useful operational data.

For large files, stream the response rather than loading it all into memory. Where servers support byte ranges, interrupted transfers can be resumed with range requests; successful partial responses use status 206 and Content-Range. A server may instead ignore a range request and send the full representation, so your downloader still needs strict size controls. See MDN’s range request guide.

Handle redirects, sessions, and browser-triggered downloads explicitly

A download link may redirect to a signed object URL, an authentication screen, or another hostname. Preserve the entire redirect chain and reapply destination policy at every hop. Do not validate only the initial URL and then blindly follow redirects.

Authentication is also part of retrieval design. Attachment requests may require the same cookies, headers, middleware, or session state used to load the source page. Scrapy’s media pipeline documentation is a useful reminder that downloader behavior—including authentication and redirect configuration—affects file retrieval, and that redirects need explicit attention in media workflows. See Scrapy’s file pipeline documentation.

Sometimes static markup contains no usable file URL because a click triggers client-side code. That is the point to use browser automation, not the default path for every page. Playwright can wait for a download event, expose a suggested filename, and save the completed download to a selected path. Its documentation also notes that browser-context downloads are temporary unless saved. See Playwright downloads.

Keep the browser path isolated from the normal HTTP downloader. It is more expensive and has different failure modes, but it solves a real class of JavaScript-triggered downloads.

Deduplicate at two levels

URL deduplication prevents repeat work during crawling. Normalize only what your domain policy permits, then key jobs by a stable identity such as normalized final URL plus relevant request context. Be cautious: query parameters can be essential for signed or versioned downloads.

Content deduplication handles the opposite case: different URLs serving the same bytes. Store a cryptographic checksum after download and use it to connect duplicate artifacts while preserving every source-page relationship.

This distinction is reflected in Scrapy’s media pipeline design: it can avoid repeated media downloads, keep original scraped URLs, record status and checksums, and retain results in the order URLs were supplied. Those are useful properties to reproduce in a custom ingestion system. Scrapy documents the metadata model here.

A durable artifact record might contain:

source_page_url
candidate_url
final_url
redirect_chain
link_text
section_heading
downloaded_at
http_status
content_type
content_length
content_disposition
safe_filename
storage_path
sha256
processor
processing_status

Safety and access controls belong in the downloader

A service that fetches discovered URLs can become an SSRF primitive if it accepts arbitrary destinations. Restrict protocols, validate every redirect destination, block private and otherwise disallowed network targets, set connection and read timeouts, cap response sizes, and limit concurrency per host. OWASP’s SSRF Prevention Cheat Sheet recommends destination allowlists and careful host validation.

Also evaluate robots directives before crawling. The Robots Exclusion Protocol is a requested access-control convention for automated clients, not authorization to access a resource; RFC 9309 makes that distinction explicit. Follow site terms, credentials, legal requirements, and operational limits in addition to crawler directives.

What PagePith demonstrated in the supplied proof

The supplied PagePith proof shows a fetch request for the MDN anchor element page. It returned the page title, reported a content length of 31,374, and included a Markdown excerpt containing the “Try it” section and link examples labeled Website, Email, and Phone.

That is a useful first-stage result: PagePith retrieved page content in Markdown form from a documentation page about anchors, where link-bearing content is central to the topic. The proof does not demonstrate attachment discovery, HTTP metadata validation, file download, browser automation, checksum generation, or file extraction. Those operations should remain explicit stages in your ingestion architecture.

Build for traceability, not just successful downloads

The best attachment pipeline is not the one that downloads the most URLs. It is the one that can answer: where did this file come from, why was it selected, what did the server return, how was it named, and which processor handled it?

Preserve that chain from source-page context through final artifact metadata. With discovery and retrieval separated, you can improve relevance rules, enforce safety limits, add browser fallbacks, and change processors without losing provenance.

Ready to test a page-retrieval workflow on your own sources? Sign up for PagePith.

Sources

  1. <a>: The Anchor elementMDN Web Docs
  2. URL() constructorMDN Web Docs
  3. RFC 6266: Content-Disposition in HTTPIETF / RFC Editor
  4. RFC 9110: HTTP SemanticsIETF / RFC Editor
  5. HTTP range requestsMDN Web Docs
  6. DownloadsPlaywright
  7. Downloading and processing files and imagesScrapy
  8. RFC 9309: Robots Exclusion ProtocolIETF / RFC Editor