← ALL FIELD NOTES

How to Turn Scraped Data into a Versioned Contract

Treat scraper output as a versioned data product: preserve evidence, publish a stable schema, test fixtures, and plan compatible migrations for downstream services.

A scraper export is already an API

A scraper may begin as a practical way to collect titles, prices, availability, or listings. The moment another job, service, dashboard, or customer-facing workflow reads its output, that JSON becomes an interface. Consumers start depending on field names, types, null behavior, identifiers, and even undocumented formatting details.

That is why a version scraped data schema practice is valuable: treat the normalized output as a data product with an explicit contract, rather than as an incidental dump from extraction code.

Scrapy makes a useful distinction here. Its items represent structured data extracted from unstructured sources, while pipelines can clean, validate, deduplicate, and persist those items before export. In other words, extraction and publication are separate responsibilities—and the published form is the one downstream systems should depend on. Scrapy items documentation and item pipeline documentation describe those building blocks.

A robust contract answers five questions:

  1. What fields can a consumer expect?
  2. Which fields are required, nullable, or optional?
  3. What does each value mean, including units and normalization rules?
  4. How can a consumer identify the schema release used for a record?
  5. What happens when the source site or extraction logic changes?

The goal is not to guarantee that a changing website will always yield every value. The goal is to make absence, change, and migration explicit rather than surprising.

Use a record envelope, not an unlabelled payload

Avoid publishing a bare object such as:

{
  "name": "Example headset",
  "price": 79.99,
  "currency": "USD"
}

It is compact, but a consumer cannot tell which schema defined it, where it came from, when it was retrieved, or which extractor created it. Put stable business fields inside a payload and surround them with metadata:

{
  "schema_id": "https://schemas.example.com/catalog/product/1.2.0",
  "schema_version": "1.2.0",
  "record_id": "product:shop.example:sku-1842",
  "source_url": "https://shop.example/products/sku-1842",
  "retrieved_at": "2026-09-22T10:15:00Z",
  "extractor_version": "2026.09.22.3",
  "payload": {
    "name": "Example headset",
    "price": 79.99,
    "currency": "USD",
    "availability": "in_stock"
  }
}

This separation matters. schema_version says how to interpret the published record. extractor_version identifies the implementation that produced it. A selector fix or parser refactor may justify a new extractor build without changing the public contract. Conversely, renaming price to current_price changes the consumer-facing interface even if the scraper code change was small.

Keep provenance with the record. Source URL and retrieval time are a practical baseline; a source snapshot reference, extraction timestamp, and transformation identifier can be useful when debugging disputed values. The W3C PROV model formalizes the idea that data entities are generated by activities and associated with agents, and it treats document versions as distinct entities. W3C PROV Primer provides the conceptual model.

Define the payload with JSON Schema

JSON Schema gives the contract a machine-readable form. Its current published specification is Draft 2020-12. JSON Schema’s specification page also describes the Core and Validation documents.

Use $schema to state the dialect and $id as the stable, unique identity of the schema. Those keywords let validators and references resolve the intended schema consistently. JSON Schema’s schema keyword reference explains both roles.

Here is a concise payload schema:

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://schemas.example.com/catalog/product/1.2.0",
  "title": "Catalog product payload",
  "type": "object",
  "additionalProperties": false,
  "required": ["name", "currency"],
  "properties": {
    "name": {
      "type": "string",
      "minLength": 1,
      "description": "Normalized display name from the source page."
    },
    "price": {
      "type": ["number", "null"],
      "minimum": 0,
      "description": "Current listed price, or null when not published."
    },
    "currency": {
      "type": "string",
      "pattern": "^[A-Z]{3}$"
    },
    "availability": {
      "type": "string",
      "enum": ["in_stock", "out_of_stock", "unknown"]
    }
  }
}

Two details deserve careful review:

  • A property in properties is not required unless it is also listed in required.
  • additionalProperties: false rejects undeclared fields.

Both behaviors are intentional choices, not syntax trivia. The JSON Schema object reference documents required and additionalProperties. Read the object keyword reference.

For a contract intended to evolve additively, strict rejection of unknown fields can be inconvenient for consumers that validate entire records. A practical pattern is to enforce a strict schema in the producer’s validation path, while consumers parse only the fields they need or use a compatibility-aware validator. Decide the policy explicitly and document it.

Choose compatibility rules before the first breaking change

A useful default is backward compatibility: a new consumer should be able to read records created with the prior schema. Under this policy, adding an optional field is usually safe. Changing an established field’s type or meaning is not.

For example, these changes are usually compatible:

  • Add optional brand.
  • Add condition with a documented default or known missing-value behavior.
  • Expand metadata that consumers do not need to interpret for existing records.
  • Mark legacy_price_text as deprecated while retaining it.

These changes need a migration plan:

  • Change price from a JSON number to a formatted string.
  • Reinterpret price from “current price” to “starting price.”
  • Rename availability without retaining the old representation.
  • Make a previously optional field required.

Schema registry documentation commonly separates compatibility against the immediately previous schema from compatibility against every historical version. This distinction matters when consumers replay long-retained records. AWS Glue Schema Registry documentation describes backward, forward, full, and transitive compatibility modes. Similarly, Confluent’s schema evolution guidance explains why optional fields and defaults affect whether older data remains readable.

Choose the stronger, transitive-style policy if a new service may read years of archived exports. Checking only the previous release is sufficient only when older records cannot realistically reappear.

Use semantic versions for the public contract

Semantic Versioning is a good release vocabulary when applied to the published contract rather than scraper internals:

  • Patch (1.2.1): clarify descriptions, repair a schema mistake that does not change valid consumer behavior, or correct extraction while the normalized meaning stays the same.
  • Minor (1.3.0): add a backward-compatible optional field or deprecate a field.
  • Major (2.0.0): remove a field, change a type, or redefine a field’s meaning.

Semantic Versioning explicitly classifies backward-compatible additions and deprecations as minor changes, and incompatible public API changes as major changes. Semantic Versioning 2.0.0 is the reference for those rules.

Validate schemas with fixtures and consumer expectations

A valid schema file is not enough. It may accept data that is technically shaped correctly but semantically useless, such as a price in the wrong unit or an unexpected availability label.

Store fixtures alongside every schema release:

contracts/
  product/
    1.2.0.schema.json
    fixtures/
      valid-in-stock.json
      valid-price-unavailable.json
      invalid-negative-price.json
      source-markup-change.json

In CI, validate each valid fixture against the intended schema. Also validate invalid fixtures and assert that they fail for the expected reason. Include examples from real source variations: no price, a discontinued item, localized numbers, alternate markup, and duplicate candidate elements.

Then add a second layer for consumer-specific assumptions. A general schema can say availability is a string from a small enum. A checkout service may additionally require that in_stock records have a non-null price. That rule belongs in a consumer contract test or consumer-owned validation, not necessarily in every producer record.

Pact distinguishes static schemas from contracts built around concrete interactions and uses consumer expectations to verify provider behavior. Its approach is a helpful model even when scraped data moves through queues or files instead of HTTP. Pact’s introduction and its explanation of consumer-driven contracts cover that distinction.

Build predictable ingestion before source changes become incidents. Define the record envelope and first fixture set for your next scraper workflow, then sign up to explore PagePith.

Deprecate fields before removing them

Suppose early records expose a fragile source string:

"legacy_price_text": "$79.99 incl. tax"

Later, you publish normalized price, currency, and tax_included. Do not silently remove legacy_price_text from a 1.x line if downstream users may still read it.

Instead:

  1. Release a minor version that adds normalized fields.
  2. Keep legacy_price_text available and mark it deprecated in the schema description and release notes.
  3. Give consumers a stated migration window.
  4. Monitor usage where that is possible in your delivery path.
  5. Remove the field only in the next major version.

For a truly incompatible change, support parallel forms temporarily. This can mean dual-writing product.v1 and product.v2 topics, serving both representations from an adapter, or retaining the old deserializer for historical data. AWS migration guidance notes that older registry-written records can require the old deserializer until those records are no longer needed. AWS Glue migration documentation illustrates the operational reason to retain old readers during transition.

Do not use dual-write as a permanent substitute for a migration. Set exit conditions: the consumers that must move, the final date, the backfill or replay decision, and the deletion criteria for the legacy path.

Make the publishing pipeline enforce the contract

A scraper architecture can enforce this sequence:

  1. Extract raw candidates from the source page.
  2. Normalize types, currencies, labels, and identifiers.
  3. Validate the payload against its pinned schema release.
  4. Attach envelope metadata including schema ID, source URL, retrieval time, and extractor version.
  5. Quarantine failures with enough evidence to debug extraction changes.
  6. Publish only accepted records and retain fixtures from important failures.

Scrapy pipelines are designed for the transformation and validation stage between extraction and storage, making them a natural location for normalization, duplicate handling, and rejection logic. Scrapy’s item pipeline guide outlines these uses.

If a downstream service exposes records over HTTP, use OpenAPI to describe endpoints, authentication, and transport behavior, while retaining JSON Schema as the payload contract. OpenAPI 3.1’s Schema Object is based on JSON Schema Draft 2020-12 with an OpenAPI dialect. OpenAPI 3.1.1 specification documents that relationship.

An honest PagePith demonstration

The supplied PagePith proof shows a fetch of the JSON Schema specification page. The returned title was JSON Schema - Specification [#section], and the Markdown excerpt included the statement that the current version is 2020-12, along with the start of a specification-documents section.

That small result demonstrates the kind of source evidence worth retaining during schema work: a requested URL, a retrieved document title, and a readable excerpt that supports a specific implementation decision. It does not demonstrate automated schema generation, compatibility enforcement, or end-to-end contract publishing. Those capabilities should be evaluated separately in your own workflow.

Start with one contract, then make change routine

The first version does not need every possible field. Start with the values that downstream services truly need, define their semantics, publish a schema identifier, and preserve raw evidence for difficult records. Add representative fixtures before the export becomes widely consumed.

From there, make every schema pull request answer four questions: Is this additive? Can a current consumer read old records? Can an old consumer safely ignore new records or fields? Which fixture proves the intended behavior?

That discipline turns web extraction from a fragile implementation detail into a governed interface that downstream services can trust.

Ready to establish a more deliberate evidence and ingestion workflow? Sign up for PagePith.

Sources

  1. JSON Schema SpecificationJSON Schema
  2. JSON Schema: The `$schema` KeywordJSON Schema
  3. Schema Evolution and CompatibilityConfluent Documentation
  4. AWS Glue Schema RegistryAmazon Web Services
  5. Semantic Versioning 2.0.0Semantic Versioning
  6. Introduction to PactPact Foundation
  7. Scrapy ItemsScrapy Documentation
  8. Scrapy Item PipelineScrapy Documentation
Version Scraped Data Schemas Safely · PagePith