How to Store Raw and Extracted Web Data for Reprocessing
A practical storage design for preserving original web captures, versioning extracted records, and tracing every result back to its source and extractor.
Web extraction is rarely correct forever. A selector breaks when a site redesigns its page. An API field gains a new meaning. A normalization rule is fixed after analysts have already loaded the prior output. If the pipeline retained only its latest CSV, JSON, or table, the team must either accept inconsistent history or collect the web again—after the source may have changed.
The durable alternative is to store raw and extracted web data as separate products with an explicit relationship between them. The raw capture is the evidence used by an extractor. The extracted dataset is a versioned interpretation of that evidence. Lineage records make the relationship reproducible.
This design does not require keeping a separate copy of raw content for every parser iteration. Capture once, preserve it safely, then create new extracted versions as extraction logic evolves.
The three-layer model
A reprocessable architecture has three durable layers:
- Immutable raw captures: the original HTTP-level material needed to run extraction again.
- Versioned extracted outputs: normalized records optimized for analysis and downstream use.
- Capture and lineage metadata: records that connect an output to exact inputs, schemas, configurations, and code versions.
These layers answer different questions:
| Layer | Main question answered | Change policy |
|---|---|---|
| Raw capture | What did we retrieve at that time? | Never overwrite |
| Extracted output | What did extractor version X produce? | Publish a new version |
| Metadata and lineage | How was this output produced? | Append or version with each run |
The common failure mode is treating an extracted table as all three. A table may be useful for querying, but it usually cannot preserve HTTP status, headers, redirects, source payload bytes, or the parser configuration that produced each value.
Preserve the HTTP exchange, not just parsed content
Saving rendered text or a parsed DOM is helpful, but it may omit details required to debug or reproduce a result. For example, a response body alone does not show whether a redirect occurred, which content type the server returned, or which request headers were sent.
The WARC 1.1 format is a useful model for archival capture because it defines records for payload content, request and response control information, associated metadata, transformed results, and duplicate-detection events. It also defines identifiers, capture dates, target URIs, and block or payload digests.
Whether you write WARC files directly or use a different container, preserve comparable information:
capture_id: "urn:uuid:3f5e..."
captured_at: "2026-09-17T11:42:18Z"
request:
method: GET
url: "https://source.example/products/42"
headers_ref: "raw/captures/2026/09/17/3f5e/request.headers"
response:
status: 200
headers_ref: "raw/captures/2026/09/17/3f5e/response.headers"
body_ref: "raw/captures/2026/09/17/3f5e/response.warc.gz"
payload_sha256: "..."
Keep sensitive data out of a broadly readable manifest. Request headers can include credentials, cookies, or other confidential values. A production design should apply access controls, redaction rules where appropriate, and a deliberate retention policy before capture begins.
Give each capture a stable identity
A URL is not a capture identity. The same URL can produce a different response every hour, and URL normalization can also collapse distinct requests unintentionally.
Assign a generated capture_id at retrieval time. Record the target URL and timestamp as attributes, then calculate and retain a content digest. WARC includes both record identifiers and digest fields, and its revisit-record model accommodates repeated content while retaining capture relationships. That makes deduplication compatible with provenance rather than a reason to discard it. See the WARC specification for the relevant record model.
A practical distinction is:
capture_id: identifies one retrieval event.payload_digest: identifies the bytes of a payload.extractor_run_id: identifies one execution of extraction logic.dataset_versionor snapshot: identifies a published extracted result.
Do not substitute one for another. Two captures may share a payload digest but still have different retrieval dates or response metadata. One extractor run may consume many captures. One published dataset version may contain outputs from several runs.
Make raw storage durable by policy, not convention
A bucket named raw is not immutable if jobs are allowed to overwrite objects at known paths. Use object keys that include a capture ID or content-derived component, deny destructive writes where possible, and treat a successful raw upload as append-only.
Versioned object storage adds recovery protection. Amazon S3 Versioning retains multiple object variants, so accidental overwrites or deletions need not eliminate prior versions. If the retention requirement calls for stronger controls, S3 Object Lock provides write-once-read-many retention and legal-hold mechanisms for protected object versions.
Integrity deserves an independent check. Calculate a checksum as part of ingestion, keep it in the manifest, and periodically verify it against the stored object. S3 supports validating supplied checksums on upload and retrieving stored checksum information later, as described in its object integrity documentation.
The goal is not to make deletion impossible in every situation. It is to make retention intentional. Define how long raw captures must remain available for parser fixes, audits, contractual obligations, and source-data change rates. Then ensure lifecycle rules do not quietly undermine that promise.
Store extracted data for queries and change
Raw captures are optimized for fidelity; analysts generally need structured, queryable outputs. Keep those concerns separate.
For extracted records, write a normalized schema to a columnar format such as Parquet. Apache Arrow’s Parquet documentation describes Parquet as a standardized columnar format for analytical I/O. A typical output row might look like this:
{
"product_id": "42",
"name": "Example Widget",
"price_amount": 19.99,
"currency": "USD",
"source_capture_id": "urn:uuid:3f5e...",
"extractor_version": "catalog-parser@a81d9e2",
"schema_version": "product-v3"
}
Include provenance columns even when the table is wide or heavily normalized. At minimum, a downstream user should be able to find the source capture, extractor version, and output batch that produced a record.
If users need reproducible historical queries, consider a table format with snapshots. Apache Iceberg creates a new snapshot on each table write and supports time travel and rollback to valid snapshots. Its metadata-based schema evolution also supports changes such as adding, dropping, renaming, reordering, and widening fields without rewriting every existing data file.
That makes a useful contract possible: a consumer can request “the product dataset as published by snapshot N,” while an engineering team can publish a corrected parser result as snapshot N+1 rather than replacing history.
Build reprocessing into the workflow from the start. If your team needs a clearer path from captured web content to versioned extraction results, sign up for PagePith.
Treat extraction code as input data
A parser version cannot be an informal release note. It is a material input to output correctness.
For every extraction run, record:
- extractor name and semantic version;
- source revision or immutable build identifier;
- container or environment identifier, when applicable;
- configuration version, including selector and normalization rules;
- input capture IDs or a precise capture-set reference;
- schema version;
- run ID, start and completion timestamps;
- output URI, dataset version, or table snapshot.
OpenLineage’s object model separates jobs, runs, and datasets and supports details including source-code locations and versions, dataset versions, schemas, inputs, and outputs. You do not need to adopt every part of that specification to benefit from its separation of concerns: a job is the logical extractor, a run is one execution, and a dataset is an input or output artifact.
The important implementation detail is exact input membership. “This parser reads the raw bucket” is pipeline-level documentation, not lineage. Record the capture IDs, a manifest with a content digest, or a frozen query result that names the exact records processed. OpenLineage’s Lineage Dataset Facet supports explicit dataset relationships and field-level lineage, rather than requiring an assumption that every upstream input contributed to every output.
Reprocess without copying raw data
Once raw captures are immutable and addressable, reprocessing should be a new extraction run over existing capture references.
Suppose catalog-parser@a81d9e2 incorrectly interpreted an empty price as zero. The repair sequence is:
- Select the affected capture IDs or a frozen manifest.
- Run
catalog-parser@b42c107against those original captures. - Validate row counts, null rates, schema conformance, and a representative sample.
- Write a new output batch or table snapshot.
- Store exact lineage from that new result to the source capture set and extractor run.
- Publish the replacement dataset version according to your consumer contract.
Raw objects are not duplicated. The changed artifacts are the extracted outputs and their associated manifests. If a source payload has already been deduplicated safely, several capture records can still reference it while preserving their own capture metadata.
Avoid mutating prior extracted files in place. Even a correction should result in a new version. This makes comparisons straightforward and lets consumers migrate deliberately.
Design retention around recoverability
Retention cleanup can break reproducibility in two places: raw objects and extracted table history.
For raw storage, noncurrent object versions may be subject to lifecycle expiration; S3 Versioning guidance notes that lifecycle configuration can manage noncurrent versions. For extracted tables, expiring Iceberg snapshots removes them from metadata and makes them unavailable for time travel, according to the Iceberg maintenance documentation.
Before expiring either layer, ask a concrete question: can we reproduce every dataset version we still promise to support? If the answer is no, shorten the published support window, retain the needed inputs longer, or preserve a durable export and its lineage.
A sensible policy may keep raw captures for a longer window than extracted snapshots, because raw material can regenerate outputs. But this is a business and governance decision, not a universal rule. Sources with legal, privacy, or contractual constraints may require stricter deletion policies.
An honest PagePith demonstration
The supplied PagePith proof shows a fetch-tier request to the WARC Format 1.1 specification. That request returned the title “The WARC Format,” a content length of 67,580, and a Markdown excerpt beginning with the document’s discussion of collecting and managing saved web objects.
That small result illustrates the first boundary in a reprocessable design: preserve what was retrieved and associate it with the request. The proof does not establish a complete archival implementation, immutable storage configuration, automated lineage, or a reprocessing interface in PagePith. Those capabilities should be verified against the requirements of your own system.
A practical implementation checklist
Before scaling extraction volume, confirm that your pipeline can answer these questions for any output row:
- Which capture produced it?
- What URL, time, status, headers, and raw payload were observed?
- Can we verify the raw object’s checksum?
- Which code and configuration version extracted it?
- Which schema governed the output?
- Which run and dataset snapshot published it?
- Can we re-run the newer extractor against the same capture set?
- Will planned lifecycle policies preserve the required recovery window?
If any answer depends on logs that rotate, object paths that are overwritten, or a developer’s memory, make that information part of the durable data model.
A reprocessable web-data pipeline is not just a storage layout. It is a contract: raw evidence remains stable, interpretations are versioned, and every published record has a path back to its source and method of creation.
Ready to make that contract easier to operationalize? Sign up for PagePith.
Sources
- The WARC Format 1.1International Internet Preservation Consortium
- Retaining multiple versions of objects with S3 VersioningAmazon Web Services
- Locking objects with Object LockAmazon Web Services
- Checking object integrity for data uploads in Amazon S3Amazon Web Services
- Reading and Writing the Apache Parquet FormatApache Arrow
- OpenLineage Object ModelOpenLineage
- Lineage Dataset FacetOpenLineage
- Apache Iceberg MaintenanceApache Iceberg