← ALL FIELD NOTES

How to Scrape Login-Protected Websites Without Hard-Coding Credentials

A practical pattern for collecting permitted account data with OAuth, managed session state, least privilege, and clear recovery paths—without embedding passwords in code.

Authenticated scraping is not mainly a parsing problem. It is an authorization, session-lifecycle, and secret-handling problem.

If a developer copies a username and password into a Playwright script, the first request may work. The operational cost arrives later: credentials leak through repositories or build logs, MFA interrupts scheduled jobs, a password rotation breaks production, or a broad employee account grants far more access than the job needs.

A safer design keeps passwords out of application code entirely. Prefer an official API with scoped access, use an approved OAuth flow where available, and reuse browser session state only when automation is permitted and no supported API can meet the requirement.

This article covers a practical architecture for permitted login-protected collection. It does not cover bypassing MFA, CAPTCHAs, access controls, rate limits, or account restrictions.

Start with permission, not a login form

A successful login does not establish permission to collect every page the account can technically reach. Authentication establishes identity; authorization determines which resources and actions that identity may use. OWASP recommends treating those as separate controls and enforcing least privilege at the resource level. Read the OWASP authorization guidance.

Before implementing a collector, write down four boundaries:

  1. Data boundary: Which records, fields, dates, and exports are needed?
  2. Account boundary: Which dedicated service account or approved user account will access them?
  3. Operational boundary: How often may the job run, and what volume is acceptable?
  4. Failure boundary: What should happen when access expires, permissions change, or MFA is required again?

For example, a finance team may authorize a daily collection of invoices for one business unit. That does not imply permission to crawl all user profiles, download unrelated documents, or use an administrator account because it is convenient.

These constraints should become code: allowlisted origins and routes, a fixed extraction schema, bounded pagination, request pacing, and an audit trail that records job identity and outcome without storing sensitive response data unnecessarily.

Choose the access method in the right order

The best implementation is usually the one that avoids browser login automation.

1. Use an official API first

If the provider has an API, use it. API credentials are generally easier to scope, rotate, audit, and revoke than a browser session. Favor a dedicated integration identity over a shared employee login.

The collector can then use a narrowly scoped token stored in a secrets manager, request only documented endpoints, and treat 401 or 403 responses as an access-state event rather than something to work around.

2. Use OAuth when the provider supports it

OAuth 2.0 allows an application to obtain limited access to protected resources without receiving the resource owner’s password. That is exactly the separation a data-collection integration needs: the user authorizes a defined scope, while the integration receives a token rather than a reusable password. RFC 6749 describes the OAuth 2.0 authorization framework.

A sound OAuth-based job typically does the following:

  • Requests the minimum documented scopes.
  • Stores access and refresh tokens in centralized secret storage.
  • Refreshes access through the provider’s supported token flow.
  • Records the authorization grant, scopes, and expiry time.
  • Stops and asks for reauthorization when refresh fails or scopes are insufficient.

Do not convert an OAuth workflow back into a password workflow by asking users to paste passwords into a configuration file.

3. Reuse approved browser session state only when necessary

Some internal portals and legacy vendor systems have no API or OAuth integration. If the operator is authorized to automate the portal, an interactive login can establish a browser session that the job later reuses.

This is not credential-free authentication. It replaces a password in source code with session material that is still highly sensitive. Playwright’s browser context state can contain cookies, local storage, IndexedDB data, and WebAuthn credentials. See the BrowserContext storage-state API.

Treat that state file like a production credential.

A practical session-state pattern with Playwright

Separate enrollment from collection. Enrollment is an approved, human-attended process that handles the provider’s normal login and MFA experience. Collection is a non-interactive job that uses the resulting state until it expires or is revoked.

Enrollment: create state outside the application repository

A controlled setup process can:

  1. Launch a browser in a secure operator environment.
  2. Let an authorized person sign in using the provider’s normal flow.
  3. Confirm the expected account and tenant are active.
  4. Save session state to a protected location.
  5. Encrypt and store that artifact in a secrets system with a short review and rotation policy.

Never commit state.json, cookies, token exports, or browser profiles to Git. Also keep them out of ticket attachments, CI artifacts, screenshots, and debug logs.

OWASP recommends centralized secret management with access controls, auditing, and rotation rather than distributing sensitive values through code and configuration. Read the OWASP secrets-management guidance.

Collection: load state and enforce a narrow route allowlist

The collection worker should receive only a path to the decrypted state artifact at runtime. The path is not the secret; the file contents are.

import { chromium } from 'playwright';

const allowedOrigin = 'https://portal.example.test';
const statePath = process.env.PORTAL_STORAGE_STATE;

if (!statePath) {
  throw new Error('PORTAL_STORAGE_STATE is required');
}

const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ storageState: statePath });
const page = await context.newPage();

const response = await page.goto(`${allowedOrigin}/reports/invoices`, {
  waitUntil: 'networkidle',
});

if (!response || [401, 403].includes(response.status())) {
  throw new Error('Session is not authorized; stop and request re-enrollment.');
}

if (new URL(page.url()).origin !== allowedOrigin || page.url().includes('/login')) {
  throw new Error('Session expired or was redirected to login.');
}

const invoices = await page.locator('[data-report="invoices"] tr').evaluateAll(rows =>
  rows.map(row => row.innerText)
);

await browser.close();
console.log(JSON.stringify({ invoiceRows: invoices.length }));

The example intentionally does not include a password, a token, a cookie value, or logic to defeat an interactive challenge. It also checks for common failure signals instead of retrying an unauthorized request indefinitely.

In a production implementation, add an explicit allowlist for the expected route family, cap the number of pages, validate that the displayed account matches the intended integration account, and send an alert when re-enrollment is required.

Protect sessions as carefully as passwords

A session cookie or bearer token can often act as the authenticated user until it expires. Do not print Cookie headers, authorization headers, storage-state JSON, or full redirected URLs in logs.

Where you control the application issuing cookies, use HTTPS-only Secure cookies, HttpOnly to reduce JavaScript access, restrictive SameSite behavior, narrow Domain and Path values, and appropriate expiry. These attributes reduce exposure but do not remove the need to protect the session artifact itself. MDN’s secure-cookie guidance explains these controls.

For the scraper infrastructure, apply comparable operational controls:

  • Decrypt session state only in the job environment.
  • Restrict secret access to the workload identity that needs it.
  • Avoid sharing one session artifact among unrelated jobs.
  • Set explicit expiry and re-enrollment expectations.
  • Rotate or revoke the session after personnel or permission changes.
  • Redact secrets before sending errors to observability tools.

OWASP also cautions against placing session IDs, JWTs, refresh tokens, and credentials in browser localStorage or sessionStorage as a general storage strategy. Its session-management guidance explains why session identifiers require dedicated protection.

Design for expiry, MFA, and permission changes

A reliable system expects authenticated access to fail eventually. Sessions expire. Administrators reduce scopes. Providers require a fresh login. MFA prompts appear after a risk event.

Your worker should classify these as normal state transitions:

SignalWorker behavior
401 UnauthorizedStop the job, mark the session invalid, and request approved reauthorization.
403 ForbiddenStop collection for that resource and investigate whether scope or authorization changed.
Redirect to loginTreat the browser state as expired or incomplete. Do not submit stored passwords.
MFA promptPause and route to a human or provider-approved enrollment process.
Expected page structure missingSave a redacted diagnostic, then alert; the portal may have changed.

MFA is a security control, not an obstacle to automate around. A human-in-the-loop renewal process can be less convenient than a fully unattended password flow, but it respects the provider’s authentication boundary and gives administrators a chance to review continued access.

Where PagePith fits—and where it does not

It is important to be precise about the phrase authenticated web scraping API. PagePith is not currently a service for fetching pages behind a login. Its documentation says it works with public URLs and does not support login-protected pages. See the PagePith documentation.

That limitation makes PagePith a fit for the public portion of a broader collection pipeline, not for replaying your authenticated portal session. For example, after your permitted internal workflow identifies a vendor’s public documentation URL, PagePith can retrieve that public page as readable content for downstream processing.

An honest PagePith demonstration

In the supplied retrieval proof, a request to https://docs.pagepith.com/ succeeded through the fetch tier and returned readable Markdown from a public page. The extracted content included the PagePith documentation title and sections such as “Make your first request,” “Extract a brand profile,” and “Monitor a page.” That is a useful example of turning a public documentation URL into normalized, application-friendly text.

It is not evidence that PagePith can access a private portal, accept exported browser cookies, manage OAuth tokens for another service, or bypass an authentication prompt. Keep those responsibilities inside your authorized integration architecture. PagePith’s public product information describes URL retrieval through a REST API and supports public content workflows. Learn more about PagePith.

Need readable content from public URLs adjacent to your authorized workflow? Create a PagePith API key.

A production checklist

Before running a login-protected collection job, verify all of the following:

  • You have explicit permission for the account, records, cadence, and intended use.
  • An official API or OAuth grant was evaluated before browser automation.
  • The job uses a least-privilege service identity or approved account.
  • No username, password, cookie, token, or browser state is committed to source control.
  • Session state is stored as a secret, access-controlled, audited, and rotated.
  • Logs and alerts redact headers, tokens, state files, and sensitive extracted data.
  • MFA or a session challenge stops the job and triggers approved re-enrollment.
  • 401, 403, login redirects, and page-layout changes have explicit handling.
  • Collection is bounded by an allowlisted origin, route set, data schema, and run schedule.
  • Public URLs are handled separately from authenticated portal URLs.

The durable solution is not to hide a password more carefully. It is to model access deliberately: scoped grants where possible, protected short-lived session state where necessary, and a clean stop when the authorization boundary changes.

For public-page retrieval and normalized content after that boundary, sign up for PagePith.

Sources

  1. PagePith documentationPagePith
  2. PagePith product homepagePagePith
  3. The OAuth 2.0 Authorization FrameworkIETF RFC Editor
  4. Session Management Cheat SheetOWASP
  5. Secrets Management Cheat SheetOWASP
  6. BrowserContext APIMicrosoft Playwright
  7. Secure cookie configurationMDN Web Docs
  8. Authorization Cheat SheetOWASP
Scrape Login Sites Without Hard-Coded Credentials · PagePith