How to Detect and Remove Boilerplate Before Sending Web Content to an AI Model
Build a reliable web-to-AI ingestion boundary: extract main content, preserve structure, measure tokens, validate quality, and handle untrusted page text.
Boilerplate removal is an ingestion problem, not a prompt problem
If you send raw web pages directly to an AI model, you are rarely sending just the article, documentation page, or product information you intended. You are also sending navigation menus, cookie banners, subscription prompts, related-post modules, footers, social links, scripts, and often several repetitions of the same site template.
The usual name for solving this is main-content extraction. It is also called boilerplate removal, web-page cleaning, or DOM-based content extraction. The goal is simple: retain the document’s useful content and discard the surrounding clutter. The Trafilatura paper frames the task in these terms: identify the main text while excluding headers, sidebars, footers, and other unwanted material.
For an AI application, this should be a deterministic stage in your ingestion pipeline—not a request such as “ignore navigation” embedded in a model prompt. A clean boundary gives you more control over:
- Retrieval relevance and chunk quality for RAG
- Input-token measurement and cost attribution
- Preservation of headings, lists, links, tables, and code
- Debugging when extraction fails on a particular domain
- Isolation of untrusted web text before it reaches an agent or tool workflow
The right objective is not “make the page shorter.” It is to produce a representation that is useful, structured, measurable, and safe to treat as external data.
What should be removed—and what should survive
A baseline extractor should remove elements that are rarely part of the primary document:
script,style, and embedded application payloads- Site navigation, headers, footers, and breadcrumbs when they are not material to the task
- Cookie and consent dialogs
- Advertisement containers and newsletter promotions
- Share widgets, recommendation modules, and “read next” cards
- Repeated links, pagination controls, and comment modules when comments are out of scope
That list is intentionally not a list of HTML tags to delete blindly. A nav element is normally noise in an article, but a navigation-like list may be essential in API documentation. A table can be a layout artifact, or it can contain the only useful comparison on the page.
Preserve semantic structure whenever downstream tasks rely on it:
- Title, byline, date, canonical URL, and site name
- Heading hierarchy
- Paragraphs and ordered or unordered lists
- Tables, code blocks, and figures when relevant
- Link text and destinations, especially in technical documentation
This is why flattening a page into plain text too early is costly. Trafilatura’s output options include structured formats such as Markdown, JSON, XML, and HTML, while Mozilla’s Readability.js can return both processed article HTML and text content. Keep a structured cleaned version for model tasks and retain raw HTML for investigation and reprocessing.
A practical pipeline for clean model input
A production pipeline benefits from explicit stages. Each stage has a testable responsibility.
1. Fetch, classify, and retain the source
Store the URL, fetch timestamp, HTTP status, response headers where appropriate, and the raw response or raw HTML. Do not overwrite the source with the extracted result.
Keeping the original makes it possible to compare extractor versions, explain a bad result, and rerun the page with different settings. Readability modifies the DOM as it parses, and its documentation recommends cloning the document when the original must remain available. See the Readability.js README.
Before extraction, decide whether the downloaded HTML contains meaningful body content. JavaScript-heavy sites may require browser rendering first. That is a separate concern from boilerplate removal: extracting from an empty application shell will not become accurate simply because the extractor has better heuristics. Zyte distinguishes HTTP-response extraction from browser-rendered HTML, noting that rendering can improve JavaScript-heavy pages while raw HTTP is typically faster and less expensive. See its automatic extraction documentation.
2. Remove obvious non-content nodes
Perform deterministic DOM cleanup before asking a main-content algorithm to choose a candidate. Remove scripts and styles, then use tag, role, class, and ID signals for recurring chrome such as cookie, consent, newsletter, sidebar, footer, and related.
Avoid relying only on CSS-class rules. Publishers use arbitrary class names, and the same word can appear in content. Treat these rules as high-confidence reductions, not as the entire extraction strategy.
3. Score candidate content blocks
After initial cleanup, score content-bearing DOM nodes. Useful signals include:
- Visible text length
- Link density: linked text divided by total text
- Paragraph count and sentence-like text
- Heading and article-context signals
- Relative DOM position
- Repetition across pages from the same host
For example, a 2,000-character main element with 3% link density and several paragraphs is a stronger candidate than a sidebar with 250 characters and 80% linked text. Trafilatura documents a similar approach: it cleans first, then uses features including text length, link density, and position to identify content nodes. Read the extraction overview.
4. Use fallbacks rather than one universal rule
Article pages, forum threads, documentation, recipe pages, and data tables have different DOM shapes. A single extractor will fail somewhere.
A reasonable implementation tries a primary extractor, checks its result, then uses fallbacks if the output is suspiciously short or structurally incomplete. Trafilatura documents fallback extractors and separate fast, precision, and recall-oriented modes. A precision-oriented path can reduce clutter more aggressively; a recall-oriented path can retain more borderline content. Neither is universally correct, so make the choice explicit for each product workflow.
For example:
- Summarizing a news article: favor precision; navigation and related stories add little value.
- Indexing technical documentation: favor recall; preserve code, tables, admonitions, and cross-references.
- Capturing a forum discussion: identify the thread body and retain individual replies, rather than forcing the page into a single article body.
5. Normalize into a model-ready document
Convert the selected DOM subtree into Markdown or structured JSON while preserving document boundaries. A useful internal shape might look like this:
interface CleanDocument {
sourceUrl: string;
title?: string;
publishedAt?: string;
extractedMarkdown: string;
headings: string[];
links: Array<{ text: string; href: string }>;
extractionMethod: string;
rawHtmlRef: string;
quality: {
characterCount: number;
linkTextRatio: number;
requiredElementsFound: string[];
warnings: string[];
};
}
This representation allows a retriever to chunk by headings, lets a UI display source links, and preserves enough diagnostics to decide whether the result should enter an index.
Want a cleaner starting point for web-content ingestion? Create a PagePith account to evaluate its workflow against your own representative pages.
Put a quality gate before the model
Do not treat a non-empty extraction as a successful extraction. The research benchmark described in the Trafilatura paper evaluates both sides of the task: preserving desired text and discarding boilerplate. A tiny output may be clean but incomplete; a long output may retain every navigation menu.
Build a quality gate that records a pass, warning, or reject decision. Useful checks include:
- Extraction success: Did the extractor identify a plausible article or content region?
- Length bounds: Is the result above a minimum useful size and below a maximum likely to indicate template leakage?
- Title consistency: Does the extracted title roughly agree with the page title or expected URL context?
- Link-text ratio: Is most of the extracted text a collection of links?
- Duplicate text: Are nav labels, headings, or paragraphs repeated unusually often?
- Error-page detection: Does the result contain login, access-denied, consent-wall, paywall, or not-found indicators?
- Required structure: For a documentation source, did code blocks or tables survive? For a product source, did the expected specification area survive?
Route warning cases to a fallback extractor, a browser-rendered fetch, or a review queue. Keep the quality metrics with the document so you can analyze failures by domain, template, or extractor version.
Measure token reduction with the target model
Cleaning pages often reduces model input because it removes repeated chrome. But do not report savings from character counts alone. Tokenization varies by model and by text, and API billing distinguishes input tokens from other token categories. OpenAI recommends measuring tokens on representative inputs; see Understanding and counting tokens.
For every sampled page, record:
raw_html_tokens
clean_markdown_tokens
reduction = 1 - (clean_markdown_tokens / raw_html_tokens)
Then review the result alongside quality, not in isolation. A 90% reduction is not useful if the extractor dropped the article’s comparison table. Conversely, a modest reduction may be acceptable if it preserves the structure needed for retrieval.
Measure by content class—articles, docs, product pages, forums—not only as a global average. That tells you where custom rules or a different extraction path are worth the maintenance cost.
Treat extracted web text as untrusted data
Boilerplate removal improves signal quality, but it is not a complete security defense. A web page can contain text intended to manipulate a model, including instructions hidden in ordinary-looking content. OWASP identifies websites and files as sources of indirect prompt injection and recommends layered controls such as validation, sanitization, separation of untrusted content from trusted instructions, least privilege, and approval for consequential actions. See OWASP’s prompt injection guidance.
Apply that guidance directly to your pipeline:
- Label page-derived text as untrusted external content.
- Place it in a separate data field or delimited context, never in system instructions.
- Do not let extracted text authorize tool calls, credential use, or data exfiltration.
- Give browsing and agent tools only the permissions required for the task.
- Validate model outputs before actions with external side effects.
- Preserve source URLs and extracted spans so suspicious output can be traced back to its origin.
Sanitize content separately if you will render extracted HTML in a browser. Readability is an extractor, not a security sanitizer; Mozilla explicitly recommends sanitization and CSP when handling untrusted input in its README.
An honest PagePith demonstration
The supplied PagePith result provides one concrete example of a source entering an extraction workflow. For the requested ACL paper URL, PagePith returned the title “Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction,” identified the result tier as pdf, and reported a content length of 42843. Its returned Markdown excerpt begins with the conference proceedings information and then the paper title, authors, and abstract.
That is useful evidence of PDF-oriented content being represented as Markdown-like text for inspection. It is not evidence that PagePith will perfectly remove boilerplate from every HTML page, render every JavaScript application, preserve every table, or prevent prompt injection. Those outcomes require testing against your pages and enforcing the quality and security controls described above.
Choose build versus managed extraction deliberately
A DIY stack built around DOM parsing and an extractor such as Readability or Trafilatura provides control over rules, observability, and storage. It is often the right choice when your source set is limited or your content types are known.
Managed services can reduce operational work around rendering, crawling, and source variation. For example, Firecrawl documents capabilities around Markdown and HTML output, raw HTML, links, screenshots, JavaScript actions, tag selection, crawling, and PDF parsing in its advanced scraping guide. Diffbot documents rendered-page classification and page-type extraction in its Extract API documentation. Evaluate these tools on your own corpus rather than assuming a vendor’s output format is sufficient for your retrieval or security requirements.
The durable architecture is the same in either case: retain the source, extract deterministically, preserve structure, measure tokens, quality-gate the output, and pass the model only validated external content.
Start with a representative test set
Create a small corpus from the pages your application actually needs: long articles, dense documentation, pages with tables and code, JavaScript-rendered routes, error pages, and pages with heavy promotional chrome. For each, define what must remain and what must disappear. Run extraction before and after every significant rule or library change.
That turns boilerplate removal from an opaque preprocessing step into an observable component of your AI system.
Ready to test a content-ingestion workflow on your own sources? Sign up for PagePith.
Sources
- Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and ExtractionAssociation for Computational Linguistics
- How extraction works — Trafilatura documentationTrafilatura documentation
- Readability.js READMEMozilla
- Understanding and counting tokensOpenAI
- OWASP Top 10 for LLM Applications — LLM01:2025 Prompt InjectionOWASP Gen AI Security Project
- Firecrawl advanced scraping guideFirecrawl
- Diffbot Extract APIDiffbot
- Zyte API automatic extractionZyte