← ALL FIELD NOTES

How to Extract Event Schedules When Dates, Times, and Locations Are Split Across the Page

Build reliable calendar-ready event records with a two-pass workflow that joins listing cards, detail pages, structured data, and API responses without mixing events.

Event pages rarely present a clean, complete record in one DOM node. A conference listing may show a session title and day, while the detail page contains the start time. A venue page may put the address in a shared footer, and a JavaScript request may carry the cancellation status. If a scraper treats every visible date, time, and location as interchangeable page-level text, it can create calendar records that look valid but describe events that never existed.

To extract event schedules from websites reliably, treat schedule extraction as a record-linking problem rather than a selector-writing problem. First discover event candidates and their identities. Then resolve each candidate into a complete, traceable record using scoped page content, structured data, detail pages, and—when appropriate—network responses.

Why schedule fields get mixed up

Consider a listing with several event cards:

  • The card title links to an event page.
  • The card shows Thursday, Oct 16 but no hour.
  • A badge beside the card says Sold out.
  • The detail page supplies 7:30 PM and a venue.
  • A page-level banner says All times local.

A broad selector such as .date, .time, or .location has no way to prove which value belongs to which title. It can accidentally join the first date in the listing with the second event's time and the footer address. The output may pass a basic schema check while being semantically wrong.

The solution is to establish an event boundary before joining fields. Schema.org event examples place the name, time, location, address, and offers under an event-specific wrapper, reinforcing a useful extraction rule: collect fields relative to the closest shared event container, not from global page matches. Schema.org also defines event-level properties for dates, location, status, previous start dates, schedules, and offers. Schema.org’s Event type is therefore a practical vocabulary for deciding what your normalized record should preserve.

Use a two-pass extraction pipeline

A robust pipeline separates inexpensive discovery from deeper resolution.

Pass 1: discover event candidates

The first pass should answer only these questions:

  1. What event-like entities are present?
  2. What is each candidate's best identifier?
  3. Is there a canonical or detail URL?
  4. Which source fields are already available?

Start with semantic data where it exists:

  • JSON-LD scripts containing Event objects
  • Microdata or RDFa marked up as events
  • <time datetime="..."> elements
  • Stable event-card containers and links
  • Sitemap URLs that appear to be event detail pages

Machine-readable values deserve priority over display strings. The HTML <time> element can expose an unambiguous datetime value even when the rendered label is abbreviated or localized. MDN’s documentation explains that the attribute is intended to represent the machine-readable date or time.

For every discovered candidate, save evidence rather than immediately producing a final event. A discovery record might look like this:

{
  "candidate_key": "listing:https://example.invalid/events#card-12",
  "title_raw": "Autumn Chamber Series",
  "detail_url": "https://example.invalid/events/autumn-chamber-series",
  "date_raw": "Thursday, Oct 16",
  "time_raw": null,
  "container_selector": "article.event-card:nth-of-type(12)",
  "source_url": "https://example.invalid/events"
}

The candidate_key is not necessarily the final identity. It is an audit handle that says where the candidate came from. Avoid using a normalized title and date as the sole identifier: title edits, recurring runs, and reschedules make that approach unstable.

Pass 2: resolve and join fields

In the second pass, visit the candidate’s detail URL when one exists. Google’s event guidance recommends a unique URL and event markup on a leaf page for each event, rather than treating a multi-event listing as the canonical event page. That makes listing pages useful for discovery and detail pages useful for completing or verifying records. See Google’s event structured-data guidance.

Resolve each field using a deliberate precedence order. A reasonable order is:

  1. Event-specific structured data on the detail page
  2. Event-specific machine-readable HTML, including time[datetime]
  3. Event-specific visible content within the detail-page main content
  4. Event-card content scoped to the matching listing container
  5. Clearly documented page-level context, such as a timezone note

Do not let a lower-confidence value overwrite a higher-confidence value without retaining both values and the reason for the choice.

For example, if a detail page provides an offset-aware startDate, prefer it to a listing’s text label. But retain the label in provenance data; it is valuable when diagnosing upstream changes.

Establish identity before joining information

A canonical detail URL is often the strongest practical event key. If there is no canonical URL, derive an internal key from stable attributes such as a site-provided event ID, source URL, and a stable DOM or API identifier. Keep the key separate from mutable schedule fields.

This distinction matters when the schedule changes. An event that moves from Friday to Saturday should usually remain the same event record with a changed start, not become a new event solely because the date changed.

At resolution time, join only data that has a defensible connection to the candidate:

  • Same container: Fields share a parent event card or event wrapper.
  • Same detail page: Fields occur in the primary content of the candidate’s canonical URL.
  • Same structured object: Properties belong to the same JSON-LD or Microdata event object.
  • Explicit reference: An API object includes the same event ID, slug, or canonical URL.

If a venue address is page-wide and the page lists multiple events, record it as a lower-confidence inherited value—or leave the venue incomplete. Guessing is more damaging than returning a validation warning.

Normalize dates without inventing time

Dates and times have different precision levels. Your data model should preserve those levels instead of coercing everything into a timestamp.

At minimum, support these states:

StateExampleCorrect treatment
Date only2026-10-16Keep as a date-only event
Local time, no offset2026-10-16 19:30Preserve local time and mark timezone as unresolved or inferred
Offset-aware time2026-10-16T19:30:00-04:00Preserve the supplied offset
All-dayFestival runs Oct 16–18Model as all-day, not midnight UTC
Unknown hourDoors open time TBDKeep the date and missing-time state

Google explicitly distinguishes date-only events, known local times with UTC offsets, and events with an unknown start hour. Its guidance warns against representing an all-day or unknown-time event as midnight UTC. Review the date and timezone guidance here.

A normalized record can make this explicit:

{
  "event_id": "source:evt_4182",
  "name": "Autumn Chamber Series",
  "start": {
    "raw": "2026-10-16T19:30:00-04:00",
    "value": "2026-10-16T19:30:00-04:00",
    "precision": "datetime",
    "timezone_source": "source_offset"
  },
  "location": {
    "name": "Main Hall",
    "address": "...",
    "confidence": "high"
  },
  "status": "scheduled",
  "source_url": "https://example.invalid/events/autumn-chamber-series",
  "warnings": []
}

When the source supplies a named timezone, retain it as well as any observed offset. This is especially important for recurring events that cross daylight-saving transitions. The iCalendar specification supports timezone-referenced starts and recurrence components such as DTSTART, RRULE, RDATE, EXDATE, and RECURRENCE-ID. RFC 5545 is a useful reference when your downstream system emits or consumes calendar data.

Need a repeatable way to turn discovered web content into structured inputs for your workflow? Create a PagePith account.

Keep recurrence rules separate from occurrences

Recurring events are a common source of duplicate records. “Every Tuesday at 6 PM” is a schedule rule; “Tuesday, October 21 at 6 PM” is one occurrence. Do not flatten the first into many guessed occurrences unless your product specifically requires a bounded expansion window and can represent exceptions.

Schema.org provides eventSchedule for repeating schedules and advises against combining that recurring schedule representation with startDate and endDate for the same schedule. Meanwhile, iCalendar supports exceptions and modified instances through recurrence-specific fields. Schema.org’s Event documentation and RFC 5545 provide the conceptual basis for keeping these entities distinct.

A practical model uses:

  • A series record with title, organizer, location, timezone, and rule.
  • An occurrence record for a concrete date/time generated or published by the source.
  • An exception record that cancels or modifies a specific occurrence.

This lets a single moved session remain attached to the series rather than appearing as a duplicate.

Track cancellations and reschedules as state changes

Do not interpret a missing event card as proof of cancellation. Listings are often paginated, filtered, or reordered. Prefer explicit source signals such as an event status, cancellation label, or detail-page update.

For a cancellation, retain the event’s identity, last known time, and location, then update status. For a reschedule, retain the old start as history and set the new start as current. Google’s guidance similarly advises keeping identifying details for cancelled events and using the previous start date when an event is rescheduled. Its event documentation describes both patterns.

Useful change fields include:

{
  "status": "rescheduled",
  "previous_start": "2026-10-16T19:30:00-04:00",
  "start": "2026-10-17T19:30:00-04:00",
  "observed_at": "2026-09-28T09:00:00Z"
}

Store source snapshots or field-level raw values when possible. They allow you to distinguish a genuine publisher change from a parser regression.

Fall back from HTML carefully

Some sites render an empty event shell and populate the schedule after JavaScript runs. In that case, use a browser session to wait for the relevant UI state, then inspect structured scripts, DOM content, and network activity. Playwright documents browser-page loading and interaction patterns in its Pages guide.

Before building intricate rendered-DOM selectors, inspect Fetch/XHR traffic. Chrome DevTools can filter network activity by Fetch/XHR and expose request payloads and response bodies, which can reveal the structured response responsible for dates, venues, pagination, or availability. Chrome’s Network reference covers those inspection features.

An API response is not automatically a license for unrestricted collection. Use site-provided routes responsibly, apply the access rules and terms relevant to your use case, and minimize load. The Robots Exclusion Protocol defines how crawler rules are published in robots.txt, while also noting that those rules are not an authorization system. RFC 9309 explains that distinction. Sitemaps can also reduce brittle pagination work by supplying a URL inventory through loc entries; see the Sitemaps protocol.

Validate records before publishing them

Validation should test relationships, not just individual formats:

  • A resolved event has a title and source URL.
  • A start is either date-only, local-naive, or offset-aware—never an invented timestamp.
  • An end does not precede a known start.
  • A location has a provenance level: event-specific, inherited, or missing.
  • A cancellation retains its identity and last known schedule fields.
  • A reschedule has both the current and previous start when the source exposes both.
  • Recurrence rules and individual occurrences are not silently merged.
  • Every final field can be traced to a URL and raw source value.

Treat warnings as product data, not merely logs. A record with timezone_unresolved or location_inherited_from_page is safer to consume than one where uncertainty was hidden by a default.

A limited PagePith demonstration

The supplied proof shows PagePith performing a fetch request for Schema.org’s Event page. The returned result identifies the page as Event - Schema.org Type, reports a content length of 139,765, and includes markdown describing an event as something that happens at a time and location, along with the beginning of its property table.

That demonstrates successful retrieval of a source page containing event vocabulary. It does not by itself demonstrate automatic event extraction, schedule joining, browser rendering, API discovery, or calendar normalization. Those steps still require an extraction design like the two-pass workflow described above and validation against the target site’s actual structure.

Build for traceability, not just completeness

The best event extractor is not the one that fills every field. It is the one that can explain where every value came from, preserve what it does not know, and update records without losing their history. Begin with structured event data, use listing pages to discover and detail pages to resolve, keep joins scoped to an event identity, and model time precision, recurrence, and status explicitly.

When dates, times, and locations are spread across a site, that discipline is what turns fragile page text into calendar-ready data.

Get started with PagePith.

Sources

  1. Event - Schema.org TypeSchema.org
  2. Event Structured DataGoogle Search Central
  3. <time>: The Time (Date) elementMDN Web Docs
  4. RFC 5545: Internet Calendaring and Scheduling Core Object SpecificationInternet Engineering Task Force
  5. Pages | PlaywrightMicrosoft Playwright
  6. Network features referenceChrome for Developers
  7. RFC 9309: Robots Exclusion ProtocolInternet Engineering Task Force
  8. Sitemaps XML formatSitemaps.org
Extract Event Schedules Reliably · PagePith