How to Extract Search Results from Public Directories and Catalogs
Build a resilient, API-first pipeline for collecting, enriching, normalizing, and monitoring records from public directories and catalogs.
Public directories, catalogs, job boards, grant listings, and searchable databases often look simple in a browser: enter a query, read a list, open a record. Production ingestion is less simple. Results may be served by an API, rendered after JavaScript runs, split across cursor-based pages, or repeated under multiple URLs.
A durable approach is to treat directory extraction as a small data-ingestion system rather than a browser-copying task. The goal is not merely to collect text. It is to produce attributable, validated records that can be refreshed safely when the source changes.
This guide outlines a practical workflow to extract data from public directories while minimizing fragility and unnecessary load.
1. Start with the least fragile access method
Before writing selectors, determine how the directory delivers its data. Prefer these options in order:
- An official API or downloadable machine-readable feed.
- A documented endpoint used by the site.
- A JSON response observed while using the public search interface.
- Server-rendered HTML result and detail pages.
- Browser automation for interfaces that genuinely require client-side interaction.
An API is usually the best contract because its fields, authentication model, and pagination behavior can be explicit. OpenAPI exists to describe HTTP APIs in a language-independent format, which can make endpoints and schemas easier to inspect and implement against. See the OpenAPI Specification.
Do not assume that a JSON request observed in a browser is a permanent or supported public interface. Evaluate its terms, authentication requirements, request volume expectations, and stability. If a documented API exists, use that instead.
A source discovery checklist
For each source, record the following before building an extractor:
- Entry URL and the exact search parameters used
- Whether results are present in the initial HTML response
- Any documented API or downloadable export
- Result-page fields, detail-page fields, and source identifiers
- Pagination mechanism: cursor, offset, page link, infinite scroll, or a
Linkheader - Whether a stable canonical URL is supplied
- Rate-limit responses, request headers, and cache validators
- Source-specific terms and collection constraints
This short reconnaissance step prevents an implementation from being tied to a visual layout that was never intended as a data interface.
2. Model results and detail pages as separate stages
Directories commonly expose two levels of information:
- Result pages provide a compact list: title, organization, category, location, status, and a link.
- Detail pages hold richer fields: descriptions, contacts, eligibility, dates, classifications, documents, or metadata.
Keep those operations separate. In the first stage, traverse result pages and emit a small discovery record. In the second stage, fetch and enrich only the records that need detail fields. The Scrapy tutorial illustrates this general pattern of extracting from listing pages, following detail links, and following pagination links independently: Scrapy Tutorial.
A compact discovery schema might look like this:
{
"source_name": "example-directory",
"source_id": "A-1042",
"source_url": "https://directory.example/records/A-1042",
"result_title": "Community Technology Grant",
"observed_at": "2026-09-24T12:30:00Z",
"search_context": {
"query": "technology",
"region": "north"
}
}
The detail-stage record can add normalized fields while preserving source evidence:
{
"source_name": "example-directory",
"source_id": "A-1042",
"canonical_url": "https://directory.example/records/A-1042",
"title": "Community Technology Grant",
"organization_name": "Example Foundation",
"location": { "region": "North" },
"description": "...",
"fetched_at": "2026-09-24T12:31:12Z",
"raw_reference": "object-storage://raw/example-directory/A-1042.json"
}
The important design choice is to retain the returned source identifier and URL. Do not derive identity by parsing an identifier from a URL when the source explicitly provides an ID. GitHub's API guidance similarly recommends relying on explicitly returned fields rather than extracting identifiers from URLs: API best practices.
3. Follow pagination supplied by the server
Pagination is a frequent cause of incomplete exports. A UI showing “Page 1” does not establish that all requests support ?page=2. A directory may use opaque cursors, session-specific filters, a Link header, or an XHR request made when the user scrolls.
Treat continuation data as part of the response contract. For HTTP APIs, inspect headers and JSON bodies for a next-page URL or cursor. GitHub documents Link headers with relationships such as next, prev, first, and last, and recommends following the server-provided URLs rather than assembling pagination requests yourself: Using pagination in the REST API.
A generic loop should make the continuation explicit:
next_url = initial_url
seen_page_urls = set()
while next_url and next_url not in seen_page_urls:
seen_page_urls.add(next_url)
response = get(next_url)
for item in parse_result_items(response):
emit_discovery_record(item)
next_url = response_next_link(response)
The seen_page_urls guard is intentionally boring. It protects against loops caused by malformed links or changing server state. Add a maximum-page and maximum-record limit as a second safety boundary, especially while developing a new connector.
For cursor pagination, preserve the original query and use the cursor exactly as returned. Do not sort, decode, or invent cursor values.
4. Extract from stable structure, not presentation details
When HTML is the appropriate source, selectors should reflect semantic structure. Prefer:
- A result-card container with a meaningful class or role
- A link's
hrefas the detail record reference data-*attributes that identify record IDs or state- Labels, headings, and scoped regions around a field
- Structured metadata embedded by the publisher, when available
Avoid selectors such as div:nth-child(3) > span:nth-child(2). They encode layout rather than meaning and tend to fail when a banner, badge, or accessibility element is added.
HTML data-* attributes are standardized hooks for associating application data with an element, and CSS selectors can target them directly. MDN's data attribute guide is a useful reference when inspecting a result page.
For example, this is easier to reason about than a positional selector:
for card in response.css('[data-record-id]'):
yield {
"source_id": card.attrib["data-record-id"],
"title": card.css('[data-field="title"]::text').get(default="").strip(),
"source_url": response.urljoin(
card.css('a[data-field="detail-link"]::attr(href)').get()
)
}
Selectors can still fail. Make failure observable by recording parsing warnings, the extractor version, and a small sample of source references. An empty result set from a normally populated query should be an alertable event, not a successful run.
5. Use a browser only when it reveals a necessary contract
Some directory pages return an almost empty document and populate results through XHR or fetch. In those cases, browser automation can help distinguish two situations:
- The browser interaction is genuinely required to obtain each record.
- The browser is only a way to discover a repeatable JSON request that is suitable to use directly.
Playwright can observe page network activity, including XHR and fetch requests, which is useful during source analysis: Playwright network documentation.
Use that capability carefully. Validate whether the request is documented or appropriate for your use case before depending on it. If direct requests are permitted and sufficient, they are generally easier to retry, cache, inspect, and operate than a full browser fleet. If an interaction remains necessary, automate the narrowest flow possible and keep the browser stage isolated from parsing and normalization.
Build connectors with a clearer evidence trail. PagePith can help you inspect fetched page content during source discovery. Create an account to try it in your workflow.
6. Normalize and deduplicate before data leaves the pipeline
Do not export raw extracted fields and postpone all quality work. Field cleanup, validation, duplicate checks, and persistence belong in the ingestion pipeline. Scrapy's pipeline documentation describes those same responsibilities, including duplicate filtering and storage: Item Pipeline.
A practical normalization sequence is:
- Trim whitespace and collapse repeated internal spaces.
- Convert missing, empty, and placeholder values into one null representation.
- Parse dates into a single standard format while retaining the source value when parsing is uncertain.
- Normalize URLs, including relative links resolved against the response URL.
- Standardize controlled vocabulary values only when the mapping is documented and reversible.
- Validate required fields and record rejects with a reason.
For deduplication, use a hierarchy of keys:
- Best:
(source_name, source_id) - Next:
(source_name, canonical_url) - Last resort: a deterministic fingerprint of normalized identifying fields
Do not use title alone as an identity key. Catalogs often contain identically named branches, repeated programs, translated listings, or annual editions.
Keep a distinction between a duplicate and an update. A stable source ID with a changed content hash is likely an update. A different URL with the same source ID is often an alias or redirect. Persisting raw payload references, fetch timestamps, and extractor versions makes that decision auditable later.
7. Throttle, retry deliberately, and support incremental refreshes
Extraction must be safe for both the source and your own workers. Respect explicit limits and avoid treating non-success responses as invitations to increase concurrency.
HTTP defines status code 429 for too many requests, and a Retry-After header can communicate when another request may be made. See RFC 9110. Your retry behavior should therefore:
- Honor
Retry-Afterwhen present. - Use bounded exponential backoff with jitter for transient failures.
- Limit concurrent requests by host.
- Stop or slow the run when error rates rise.
- Separate permanent validation failures from retryable transport failures.
Latency-aware throttling is also useful for crawlers. Scrapy's AutoThrottle documentation describes adjusting delays based on response latency and avoiding delay reductions based on non-200 responses: AutoThrottle extension.
For scheduled refreshes, avoid downloading and reprocessing every detail page if the source offers validators. Conditional requests using ETags or Last-Modified can return 304 Not Modified, reducing transfer and unnecessary work; GitHub documents this pattern in its API best practices.
8. Preserve provenance and apply collection boundaries
A clean record without provenance becomes difficult to trust. At minimum, store:
source_namesource_idsource_urlandcanonical_urlwhen known- Fetch timestamp and response status
- Search query or category path that discovered the record
- Extractor version
- A raw-payload reference or content hash
This metadata lets downstream users answer basic questions: Where did this field come from? Which run saw it? Did a source change, or did the parser change?
Also decide what should not be collected. Review the directory's applicable rules, collect only fields needed for the stated purpose, protect retained data, and set retention and deletion behavior before a large backfill. Robots directives can be a useful traffic-policy input, but they are not authentication and do not replace a source-specific permissions review.
An honest PagePith fetch example
The supplied PagePith proof shows a fetch-tier request for https://spec.openapis.org/oas/v3.0.3. It returned the title “OpenAPI Specification v3.0.3”, a content length of 160,281, and a Markdown excerpt containing specification text and a table.
That is a useful source-discovery outcome: a fetch can expose readable Markdown-like content from a standards page, making it easier to inspect document structure before designing an extractor. It does not demonstrate pagination handling, browser automation, field-level extraction, scheduled monitoring, or successful ingestion into a database. Those capabilities need separate validation against the specific directory you intend to collect.
Make every connector repeatable
A reliable public-directory connector has a small, explicit contract: how records are discovered, how details are fetched, how pages continue, how identity is assigned, and how failures are handled. Start with the official interface where possible. Keep discovery separate from enrichment. Follow server-provided pagination. Normalize before storage, retain evidence, and slow down when the source asks you to.
That approach produces a pipeline that is easier to refresh, debug, and adapt when a catalog changes its interface.
Ready to inspect a source page as part of your ingestion design? Sign up for PagePith.
Sources
- OpenAPI Specification v3.0.3OpenAPI Initiative
- Using pagination in the REST APIGitHub Docs
- Scrapy TutorialScrapy Documentation
- Item PipelineScrapy Documentation
- Use data attributesMDN Web Docs
- NetworkPlaywright Documentation
- RFC 9110: HTTP SemanticsIETF
- AutoThrottle extensionScrapy Documentation