← ALL FIELD NOTES

How to Extract Changing Web Tables Without Corrupting Data

Build a resilient extraction pipeline for web tables whose columns, merged cells, headers, and responsive layouts change over time.

A table scraper often fails quietly. The request still succeeds, the selector still finds rows, and the job still writes records. But after a publisher merges two headings, reorders columns, or serves a mobile-specific table, a value can land in the wrong downstream field.

That is the failure mode to design against when you extract changing web tables. The core problem is not selecting td elements. It is converting an evolving visual grid into a stable, auditable record model.

A durable pipeline treats every table as two related artifacts:

  1. a versioned representation of table geometry and semantics, and
  2. a normalized set of records emitted only when that representation passes validation.

This approach makes a layout change observable before it becomes data corruption.

Start with the table model, not the first row

It is tempting to assume that the first tr is a header and every later row contains values. That assumption fails on tables with captions, grouped headers, subtotal rows, multiple body sections, or footer rows.

The HTML table model distinguishes components such as caption, colgroup, thead, tbody, tr, and tfoot. It also defines how cells participate in a coordinate grid. Read the relevant sections before making assumptions about a source's structure: HTML Standard — Tables.

For each extraction run, preserve enough evidence to re-evaluate those assumptions later:

  • Request URL, final URL, retrieval timestamp, status, and response headers
  • Raw response body when available
  • Rendered DOM snapshot when browser rendering is used
  • The selected table's serialized markup or a stable fragment
  • Extraction configuration, including viewport and visibility policy
  • The inferred schema fingerprint and validation results

The source and rendered snapshots answer different questions. A server response may contain an incomplete shell while client-side code adds rows later. Conversely, a source can include alternate desktop and mobile variants, only one of which is visible in a particular rendering context.

Use a staged acquisition strategy

Begin with the least complex representation, then escalate only when the required table is not present or not complete.

1. Fetch the HTTP response

Fetch the page and identify candidate tables from the response body. Store validators such as ETag and Last-Modified when the server provides them. On later runs, conditional requests can avoid downloading and reprocessing unchanged representations; HTTP defines validators and 304 Not Modified behavior in RFC 9110.

A simplified acquisition record might look like this:

snapshot = {
    "url": final_url,
    "retrieved_at": now_utc(),
    "etag": response.headers.get("etag"),
    "last_modified": response.headers.get("last-modified"),
    "representation": "response_html",
    "body": response.text,
}

A 304 is useful for scheduling and cost control, but it is not a substitute for schema history. A new representation may retain the same logical schema, and a changed representation may be only a cosmetic change. Compare normalized structure when content is new.

2. Render when the initial response is insufficient

If the expected rows, columns, or table are absent from the response, use a browser stage. Playwright supports navigation and execution in the page context, while its locators resolve against the current DOM after re-renders. See Pages and Locators.

At this stage, record the viewport and the wait condition. Avoid a vague fixed delay such as “wait three seconds.” Wait for an observable condition: a table locator has a minimum row count, a loading indicator disappears, or a known API-driven status settles.

3. Decide explicitly how to handle hidden variants

Do not let CSS visibility decide your data policy by accident. A page can carry more than one table representation, and hidden structures may be intended for another viewport. pandas.read_html exposes a displayed_only option, illustrating that visibility changes parsing results; its documentation also cautions that HTML cleanup may still be necessary. See pandas.read_html. CSS display changes can also affect layout and accessibility-tree representation, as documented by MDN.

Define a policy such as:

  • extract the canonical desktop table at a fixed desktop viewport;
  • extract all variants and select one by explicit attributes or surrounding labels; or
  • prefer a source-provided data endpoint when that endpoint is available and appropriate for the use case.

Persist the policy with the snapshot. It is part of the extractor's contract.

Normalize geometry before interpreting columns

The most important implementation detail is an explicit rectangular grid. colspan and rowspan mean a cell occupies multiple coordinates; the HTML model requires that cells do not overlap. rowspan="0" has special row-group behavior, spanning remaining rows in that group. These details are defined in the HTML table specification.

Do not map a row's raw cell list directly to column indexes. First materialize placement.

Consider a visual structure with a two-level header:

                 2025                 2026
Item       Planned   Actual      Planned   Actual
Copper          10       11           12       13

The first heading row contains cells that span multiple columns. After expansion, produce a grid where every coordinate is occupied deliberately:

Item | 2025 / Planned | 2025 / Actual | 2026 / Planned | 2026 / Actual

A useful span-expansion algorithm maintains an occupancy map for each output row:

  1. Traverse source rows in section order.
  2. Before placing a source cell, advance to the next unoccupied output column.
  3. Parse its effective column span and row span.
  4. Mark every covered coordinate as occupied by that cell.
  5. Reject a collision rather than overwriting it.
  6. Continue until all source cells are placed.
  7. Pad only when your policy permits it; otherwise quarantine uneven rows.

The intermediate representation should preserve provenance, not merely text:

GridCell(
    text="Actual",
    source_row=1,
    source_cell=2,
    grid_row=1,
    grid_col=2,
    row_span=1,
    col_span=1,
    tag="th",
    attributes={"scope": "col"},
)

This gives you a way to explain why a downstream field received a value. It also makes a geometry diff possible: compare occupancy, spans, section membership, and header relationships rather than comparing only rendered strings.

Derive stable identities from header paths

Column position is an implementation detail. A better column identity is a normalized header path plus selected semantic metadata.

For the example above, the values might become:

item
2025.planned
2025.actual
2026.planned
2026.actual

Construct paths from header cells that apply to each data-grid coordinate. Prefer explicit semantics when provided:

  • scope="col" and scope="colgroup" identify column-oriented headers.
  • scope="row" and scope="rowgroup" identify row-oriented headers.
  • headers can explicitly point to relevant header-cell IDs.

The W3C describes these associations, including cases where complex tables need explicit relationships, in Technique H63. The HTML specification requires headers references to resolve to participating header cells in the same table, making a valid explicit association stronger evidence than visual proximity alone: HTML Standard — Tables.

A practical precedence order is:

  1. valid explicit headers relationships;
  2. applicable scoped headers;
  3. headers inferred from the expanded grid and table sections;
  4. positional inference, marked as lower confidence.

Normalize label text carefully: trim whitespace, collapse repeated spaces, retain the original label, and use a source-specific alias map only when you can document it. Do not silently turn Actual into Forecast because a matching rule happened to fit.

Build extraction jobs with evidence, not just output. PagePith can fetch a page into Markdown for inspection before you design table-specific normalization. Create an account.

Version schemas and emit changes as events

Once each column has a semantic identity, create a canonical schema object and hash it. Include more than names:

schema = {
    "table_key": "annual-materials-output",
    "columns": [
        {"key": "item", "path": ["Item"], "type": "string"},
        {"key": "2025.planned", "path": ["2025", "Planned"], "type": "number"},
        {"key": "2025.actual", "path": ["2025", "Actual"], "type": "number"},
    ],
    "header_depth": 2,
    "geometry": {"width": 3, "header_rows": 2},
}

When a new snapshot arrives, compare it to the last accepted schema and emit a structured event. Typical event types include:

  • column_added
  • column_removed
  • column_reordered
  • header_label_changed
  • header_path_changed
  • span_geometry_changed
  • table_variant_changed
  • data_row_width_changed

Not every event should block extraction. A pure reorder may be safe if identities are unchanged. A header-path change from 2025 / Actual to 2025 / Estimate is semantically material and should normally require review. The goal is to make coercion an explicit decision rather than an accidental side effect of array indexes.

Validate before records reach downstream systems

Parsing is not validation. Keep a quarantine path for suspicious snapshots and rows, with enough saved evidence for a developer to diagnose the issue.

Useful invariants include:

  • Every expanded row has the expected rectangular width.
  • No cell placement collides with an already occupied coordinate.
  • Every emitted data field has a complete header path or an approved exception.
  • headers references, when present, resolve within the selected table.
  • Data rows do not unexpectedly become header-only or empty separator rows.
  • Numeric and date conversion failures are counted and compared with prior runs.
  • The selected table still matches its identity signals, such as caption, nearby heading, or expected header keys.

A row that violates an invariant should not be force-fit into a prior schema. Store its raw cells, expanded cells, inferred paths, and failure reason. This is especially important when a publisher introduces a subtotal row that resembles data, or a responsive rendering replaces cells with labels and values inside a different layout.

Where parsing libraries fit

pandas.read_html can be productive for quickly finding and loading conventional tables. It supports matching and attribute filtering, header and index choices, link extraction, visibility behavior, and attempts to handle rowspan and colspan. But its own documentation notes that cleanup can be required for HTML idiosyncrasies: pandas.read_html.

Use it as a parsing component, not as your schema contract. For stable production extraction, keep your own snapshot, geometry normalization, header-resolution, validation, and schema-diff layers around whichever parser you choose.

A small, honest PagePith demonstration

The supplied PagePith proof shows a fetch of the HTML Standard's tables page. The recorded result has the title HTML Standard, uses the fetch tier, and reports a content length of 110451. Its Markdown excerpt contains the table-related table of contents, including entries for tbody, thead, colgroup, and related table elements.

That demonstration supports a narrow but useful workflow: fetch a standards or source page, inspect its Markdown-oriented result, and use that inspection to guide extraction design. It does not demonstrate browser rendering, automatic table selection, merged-cell expansion, schema inference, or change monitoring. Those capabilities should be evaluated separately in your own pipeline.

Design for reviewable change

A changing table is not inherently unreliable. It is simply a source whose presentation and schema need separate handling. Preserve representations, expand spans into coordinates, derive semantic header paths, validate every run, and treat schema drift as an event.

With those boundaries in place, selector breakage becomes only one class of failure—not the mechanism that silently rewrites the meaning of your data.

Ready to inspect web pages before turning them into extraction rules? Sign up for PagePith.

Sources

  1. HTML Standard — TablesWHATWG
  2. Using the scope attribute to associate header cells with data cells in data tablesW3C WAI
  3. pandas.read_htmlpandas documentation
  4. PagesPlaywright
  5. LocatorsPlaywright
  6. display CSS propertyMDN Web Docs
  7. RFC 9110: HTTP SemanticsRFC Editor / IETF