← ALL FIELD NOTES

How to Let an AI Agent Compare Two Web Pages Without Losing Their Section Structure

Build a section-aware webpage comparison pipeline that renders, extracts, aligns, and diffs meaningful page regions for AI agents.

A whole-page text diff answers the least useful version of a change question: “Are these byte sequences different?” An AI agent usually needs something closer to: “The pricing copy changed under Main > Pricing > Enterprise; the navigation order also changed, but the article body did not.”

That answer requires structure to survive the pipeline. Rather than flattening each page and diffing the resulting strings, capture a structured representation of each page, align comparable regions, then calculate changes within those aligned regions.

This approach produces results that are easier for an agent to explain, filter, route, and verify.

Why flat page diffs fail agents

A raw HTML diff is too close to implementation details. A class-name change, an A/B-test wrapper, or a new analytics attribute can overwhelm a meaningful sentence edit. A whole-page text diff is cleaner, but it still loses the location and role of the text that changed.

Consider these two revisions:

Revision A
H2: Enterprise
Paragraph: Contact sales for custom pricing.

Revision B
H2: Enterprise
Paragraph: Annual plans include priority support.

A text diff can show deletion and insertion. But it cannot reliably state where the change belongs if the same sentence appears in a footer, modal, or navigation panel too.

HTML already contains useful starting signals. The HTML standard distinguishes sectioning content and heading content, while a section represents a thematic grouping that commonly has a heading. Use those signals to create an initial content map rather than treating a document as one undifferentiated text field. WHATWG’s document structure definitions and its guidance for section are useful references for this model.

The goal is not to perfectly reconstruct a human editor’s intent. The goal is a stable, inspectable representation that lets an agent associate a change with a region, heading path, source evidence, and match confidence.

Model a page as regions and sections

Start with two nested concepts:

  • Region: a broad programmatic page area, such as main, nav, header, footer, article, or aside.
  • Section: a content grouping within a region, identified by semantic elements, heading hierarchy, or both.

Semantic regions are useful alignment anchors when markup wrappers change. W3C describes landmarks and semantic elements as programmatic regions that distinguish navigation, main content, complementary content, and related areas. Its semantic-region technique is a good basis for choosing the region labels your extractor retains.

For every extracted section, preserve at least:

{
  "id": "main/pricing/enterprise",
  "region": "main",
  "headingPath": [
    { "level": 1, "text": "Pricing" },
    { "level": 2, "text": "Enterprise" }
  ],
  "sourceOrder": 12,
  "text": "Contact sales for custom pricing.",
  "links": [
    { "text": "Contact sales", "href": "/contact" }
  ],
  "warnings": []
}

The headingPath is central. It gives the agent a human-readable address, while sourceOrder remains a fallback signal when labels are missing or duplicated.

Do not “repair” a bad heading tree by silently changing levels. Heading ranks communicate organization and subsection relationships, but real pages can skip levels or use headings inconsistently. Preserve the observed rank and attach a warning such as heading-level-jump instead. That retains evidence for downstream policy decisions. W3C’s heading guidance explains why hierarchy and rank matter.

A practical extraction order

Use a layered strategy rather than relying on one tag:

  1. Identify broad landmarks and semantic regions.
  2. Walk descendants in DOM order.
  3. Create explicit sections for meaningful section and article elements.
  4. Build implied sections when headings begin a new hierarchy level.
  5. Attach paragraphs, lists, tables, and links to the nearest active section.
  6. Retain an ancestry-based fallback ID for unlabeled content.

This lets the representation remain useful even when a page uses generic div containers instead of ideal semantic markup.

Render first, then define what “text” means

For modern sites, extracting the server response alone may capture a shell rather than the page a visitor sees. Use a browser-rendering step and wait for a declared readiness condition: an important locator is visible, a known loading indicator disappears, or a suitable load state occurs.

Avoid fixed sleeps as the primary synchronization method. Playwright documents page-level waiting and cautions that timeout-based waiting is flaky in production; use observable page conditions instead. See the Playwright Page API.

Then make an explicit product decision about your comparison target:

  • Source comparison: what arrived in the initial HTML.
  • Rendered comparison: what the browser DOM contains after the readiness condition.
  • Visible-content comparison: text intended for a reader.
  • Article comparison: the primary editorial content only.

Do not treat textContent as equivalent to visible reader text. MDN notes that it includes descendant script and style text and can include hidden content; its behavior differs from style-aware innerText. MDN’s textContent reference explains the distinction.

In most agent workflows, rendered comparison plus an explicit boilerplate policy is a good baseline. Exclude scripts, styles, templates, and known page chrome where appropriate, but record that policy in the output. Otherwise, a consumer cannot tell whether a missing section was removed from the page or filtered before comparison.

Need structured page evidence for an agent workflow? Create an account to explore PagePith: sign up.

Align sections before diffing their contents

Once both pages are represented as section trees, alignment becomes the main engineering task. Do not pair sections only by array index: one inserted section would shift every subsequent match.

A pragmatic alignment score can combine several signals:

score =
  0.40 × normalized heading similarity +
  0.25 × heading-path similarity +
  0.20 × region compatibility +
  0.10 × content similarity +
  0.05 × source-order proximity

The exact weights are policy choices, not universal facts. What matters is that each match carries its score and the signals that produced it.

For example, match Main > Pricing > Enterprise across revisions primarily by its heading path. If the heading changes from “Enterprise” to “Business,” supporting signals—same main region, neighboring sections, similar body content, and similar source position—may still justify a match. Return it as modified-heading, not as one removal plus an unrelated addition.

For repeated headings such as “Overview,” use parent paths, region, and local neighbors to disambiguate. When competing candidates score similarly, return an ambiguous result rather than manufacturing certainty:

{
  "status": "ambiguous",
  "beforeSectionId": "main/features/overview",
  "candidateAfterSectionIds": [
    "main/platform/overview",
    "main/security/overview"
  ],
  "confidence": 0.48
}

A useful rule: alignment decides which units correspond; diffing decides how their content differs. Keeping those stages separate makes failures diagnosable.

Diff within the matched section

After matching, compare normalized section content. Normalize cautiously: collapse incidental whitespace, canonicalize line endings, and optionally normalize tracking parameters in URLs only if that matches your use case. Do not erase punctuation, numbers, or link targets that may be meaningful changes.

Represent changes at two levels:

  1. Low-level evidence: insertions, deletions, and unchanged spans.
  2. High-level classification: added, removed, modified, moved, reordered, or unchanged.

Diff Match and Patch exposes insertion, deletion, and equality operations, along with semantic cleanup intended to produce more readable matches. Its API documentation is a useful implementation reference.

For an agent, a structured object is more durable than a prose-only summary:

{
  "schemaVersion": "1.0",
  "changeType": "modified",
  "before": {
    "sectionId": "main/pricing/enterprise",
    "headingPath": ["Pricing", "Enterprise"],
    "text": "Contact sales for custom pricing."
  },
  "after": {
    "sectionId": "main/pricing/enterprise",
    "headingPath": ["Pricing", "Enterprise"],
    "text": "Annual plans include priority support."
  },
  "confidence": 0.94,
  "evidence": {
    "matchSignals": ["same-region", "same-heading-path"],
    "operations": ["delete", "insert"]
  }
}

An agent can turn this into a user-facing explanation, trigger a review when pricing changes, or ignore reorder-only navigation changes—all without reparsing an unstructured narrative.

Distinguish moves, additions, and template noise

A section with substantially similar content but a different source position may be a move, not a removal and addition. Detect this after initial matching: if the content and structural signals agree but source order differs materially, emit moved or reordered with both positions.

Keep navigational and template content in the representation, but give consumers controls to exclude it. A site-wide footer-link change may matter to a release-monitoring agent but not to an editorial-change agent. Region labels make this a query decision rather than a lossy extraction decision.

Article extraction can be a valuable additional view for editorial comparisons. Mozilla Readability can parse article content and metadata, but its documentation notes that parsing changes the DOM unless a clone is used and that its output is not a stable API. Treat it as an optional normalization layer, not your canonical structural representation. See Readability.js.

Version the result and preserve failures

A comparison service is an interface between extraction code and agent logic. Version its schema from the beginning. The result should distinguish:

  • a successful comparison with no meaningful changes,
  • a rendering or extraction failure,
  • a low-confidence or ambiguous alignment,
  • a policy-driven exclusion,
  • a genuine content change.

Schema validation is a practical guardrail for this boundary. JSON Schema is designed for validating structured instances and reporting validation results, so it fits versioned comparison objects well. The JSON Schema specification draft describes this validation model.

Also treat page content as untrusted. If you retain extracted HTML for review, sanitize it before any later rendering, and avoid executing target-page scripts in your result viewer. Readability explicitly warns that extraction is not sanitization. Its security guidance is directly relevant.

An honest PagePith demonstration

The supplied PagePith proof shows a fetch of the HTML Standard document that returned the title “HTML Standard”, a reported content length of 209,761, and Markdown beginning with a nested table of contents for “Semantics, structure, and APIs of HTML documents.”

That is relevant to section-aware comparison because the returned Markdown preserves visible hierarchy in the excerpt: a top-level chapter, a numbered subsection, and deeper numbered entries. It is a concrete example of structural cues surviving extraction rather than becoming one flat block of text.

The proof does not establish browser rendering, JavaScript execution, semantic landmark extraction, section alignment, or change detection. Those functions should therefore be validated independently in your own pipeline before you rely on them for agent decisions.

Build for traceability, not just summaries

The best agent-facing comparison is not the one with the cleverest one-line summary. It is the one that can answer follow-up questions: which section changed, what matched it, what text supports the conclusion, what was excluded, and how confident is the system?

Preserve regions, heading paths, source order, matching signals, and low-level diff evidence. With those pieces in place, an AI agent can compare two web pages as structured documents instead of pretending they are plain strings.

Ready to evaluate a structured extraction workflow for your own agent? Sign up for PagePith.

Sources

  1. HTML Standard: Document structure and content categoriesWHATWG
  2. HTML Standard: The section elementWHATWG
  3. Headings | Page Structure TutorialW3C Web Accessibility Initiative
  4. Using semantic HTML elements to identify regions of a pageW3C Web Accessibility Initiative
  5. Node: textContent propertyMDN Web Docs
  6. Page | Playwright APIMicrosoft Playwright
  7. Readability.jsMozilla
  8. Diff Match and Patch APIGoogle