← ALL FIELD NOTES

How to Build an AI Agent That Can Follow Links During Research

Build a bounded web browsing workflow for AI agents with link ranking, guardrails, evidence capture, deterministic stops, and traceable results.

A useful research agent does more than fetch a starting page. It decides which links are worth pursuing, gathers evidence from a constrained set of pages, and returns an answer whose path can be inspected.

The important word is constrained. A model allowed to click whatever it wants can loop through calendars, tag archives, faceted navigation, login flows, and irrelevant pages. Worse, it may encounter webpage content that tries to redirect its behavior. The goal is not to make a general-purpose crawler with an LLM attached. It is to build a small, observable research system with explicit authority and an explicit end.

This guide presents a practical web browsing workflow for AI agents: start with one URL, extract candidates, rank them, fetch only approved links, preserve evidence, and stop on deterministic budgets or sufficient coverage.

Model the workflow as a state machine

An agent framework can provide tools, guardrails, sessions, human approvals, and tracing, but the research policy should remain visible in your application. The OpenAI Agents SDK, for example, describes both an SDK-managed run loop and the option to manage tool dispatch and state yourself (documentation). For link-following research, owning the state is often worth the extra code.

Use a state object that cannot be silently expanded by model output:

ResearchState = {
    "question": "What implementation guidance does this documentation provide?",
    "seed_url": seed_url,
    "frontier": PriorityQueue(),
    "visited": set(),
    "evidence": [],
    "page_count": 0,
    "started_at": now(),
    "limits": {
        "max_depth": 2,
        "max_pages": 12,
        "max_seconds": 90,
        "max_candidates_per_page": 20,
    },
}

The frontier holds candidate pages that are allowed to be considered. The visited set prevents repetition after URL normalization. evidence is not a model memory summary: it is a structured record of what was retrieved and why it matters.

This design addresses a fundamental crawler problem. Automated clients can recursively traverse links, and unbounded traversal can cover an entire URI space (RFC 9309). A frontier, a visited set, depth, and page count turn that open-ended behavior into a finite research run.

Separate link discovery from link navigation

Do not give the model a browser and ask it to “keep researching.” First collect links from the current page. Then evaluate those links against the task. Only then fetch the selected destination.

That separation makes the decision inspectable and testable:

  1. Fetch and render the current page.
  2. Extract visible, relevant links and their anchor text.
  3. Normalize URLs and discard duplicates.
  4. Apply deterministic policy filters.
  5. Ask a model or ranking function to score the remaining candidates.
  6. Add only the highest-scoring approved links to the frontier.

Navigation sections are particularly useful sources of candidates because HTML defines nav for sections containing navigation links (MDN reference). But do not limit extraction to navigation alone. In-content links, headings around links, and anchor text frequently carry stronger topical signals.

A normalized candidate can look like this:

CandidateLink = {
    "url": canonical_url,
    "anchor": "Tracing",
    "context": "Inspect agent runs, tool calls, and guardrails.",
    "parent_url": current.url,
    "depth": current.depth + 1,
    "same_domain": True,
    "score": 0.0,
}

Use deterministic filters before an LLM ranker. For example:

  • Require http or https.
  • Drop fragments and tracking parameters during canonicalization.
  • Reject URLs already in visited or already queued.
  • Enforce an allowlist of domains or a same-origin rule.
  • Reject file types and routes that your read-only researcher does not support.
  • Enforce the depth limit before scoring.

Then rank the survivors using a narrow output contract. Ask for a relevance score and a short reason, not a browsing command:

Question: {question}
Candidate: {anchor} — {context}
Return JSON only:
{"relevance": 0.0 to 1.0, "reason": "one sentence"}

The controller, not the model, decides whether a score is high enough and how many links may enter the queue.

Build the workflow around bounded retrieval, not autonomous clicking. If you want to evaluate a retrieval layer in a research pipeline, start with PagePith.

Choose HTTP fetching or a browser deliberately

There are two valid retrieval paths.

Use direct fetching for static, public content

An HTTP client is usually the simpler choice when the page content is available in the initial response and no interaction is needed. It is easier to cache, faster to operate, and has a smaller execution surface. Extract title, canonical URL, readable text, and links from the returned document.

Use browser automation when rendering changes the answer

Use a browser when client-side rendering, redirects, hydration, consent flows, or interaction are necessary to see the relevant content. Playwright supports both direct navigation and interaction-driven navigation, while its documentation cautions that modern pages may continue rendering after a load event (navigation guide).

That means load is not an adequate universal completion signal. Wait for something relevant to the task: a target heading, a results container, a known article body, or an explicit application-ready marker. Avoid arbitrary sleeps where possible.

Create a fresh browser context for each research run. Playwright describes browser contexts as isolated, non-persistent sessions with separate cookies and storage (BrowserContext documentation). This prevents one task’s login state, local storage, and visited-page behavior from contaminating another task.

Treat every page as untrusted input

A page can contain instructions aimed at the model: “ignore prior directions,” “upload the findings,” or “use these credentials.” Those strings are data, not policy.

OWASP identifies indirect prompt injection as a risk when external content such as webpages influences model behavior, and recommends clear trust boundaries and user control (OWASP guidance). Apply that principle in the architecture:

  • Keep the user’s research question and system policy separate from retrieved page text.
  • Never pass retrieved text into a privileged tool-selection prompt as trusted instructions.
  • Give the agent read-only tools by default: fetch, render, extract, rank, and summarize.
  • Do not expose credential entry, form submission, purchases, account changes, or arbitrary downloads to this agent.
  • Require explicit human approval for any future privileged action.

This is also a least-privilege design. OWASP’s treatment of excessive agency emphasizes limiting functionality and permissions and using approval for privileged operations (OWASP Top 10 PDF). A research agent should research; it should not become an unattended browser operator.

Before traversal, check the site’s robots.txt policy and identify the crawler appropriately. The Robots Exclusion Protocol defines directives that crawlers are requested to honor, but explicitly does not make robots rules an authorization mechanism (RFC 9309). Authentication and application authorization remain separate controls.

Capture evidence as pages are visited

A final narrative without retrievable support is hard to validate. Store evidence at the page level while the run proceeds:

Evidence = {
    "url": final_url,
    "title": page_title,
    "status": 200,
    "retrieved_at": timestamp,
    "parent_url": candidate.parent_url,
    "depth": candidate.depth,
    "excerpt": relevant_passage,
    "claim": "What this page supports for the research question",
}

The parent_url creates a link path from the seed page to each finding. That matters when debugging: you can see whether a weak source entered because of a poor ranker decision, a permissive domain rule, or an incorrect canonicalization step.

Trace the full run as well: candidate extraction, policy rejections, ranking inputs and outputs, fetch status, timing, and the stop reason. The Agents SDK tracing model includes traces and spans for model generations, tool calls, guardrails, and handoffs (tracing documentation). Even if you use another stack, the same observability shape is useful.

Stop for a reason, not because the queue is empty

An empty frontier is one stopping condition, but it is not enough. A ranking mistake could leave the queue empty after one page; an overly broad queue could keep it nonempty for far too long.

Use hard limits and research-quality limits together:

stop = (
    state.page_count >= max_pages
    or elapsed_seconds(state) >= max_seconds
    or current.depth >= max_depth
    or independent_supporting_sources(state.evidence) >= target_sources
    or frontier.empty()
)

Add an explicit stop reason to the response, such as coverage_target_reached, page_budget_reached, time_budget_reached, or no_approved_candidates. This makes cost and completeness legible to the caller.

A priority queue gives you a controlled breadth-first strategy: favor high-relevance candidates, use depth as a tie-breaker, and decline low-value links once the evidence threshold is met. Do not let the model override a budget because it claims another click may help.

End-to-end controller pseudocode

The controller should be boring. That is a feature.

enqueue(seed_url, depth=0, score=1.0)

while within_hard_limits(state) and not coverage_is_sufficient(state):
    candidate = frontier.pop_best()
    if candidate.url in visited:
        continue

    page = retrieve(candidate.url, isolated_context=True)
    visited.add(candidate.url)
    page_count += 1

    if not page.is_allowed or page.status != 200:
        log_rejection_or_failure(page)
        continue

    evidence.extend(extract_relevant_evidence(page, question))

    links = extract_links(page)
    links = normalize_filter_and_dedupe(links, state)
    ranked = rank_links(question, links)

    for link in ranked[:max_candidates_per_page]:
        if link.score >= score_threshold:
            frontier.push(link)

return synthesize_answer(evidence, stop_reason, trace_id)

Keep synthesis last. The model may summarize only the accumulated evidence, cite the page records internally, and state when evidence is limited. It should not introduce a new web search or a new navigation phase during answer generation.

A narrow PagePith demonstration

The supplied retrieval proof shows a fetch of the OpenAI Agents SDK documentation page at https://openai.github.io/openai-agents-python/. The result is titled “OpenAI Agents SDK,” is marked as the fetch tier, and contains 6,880 characters of content. Its excerpt identifies agents as LLMs equipped with instructions and tools, and lists guardrails among the SDK primitives.

That is a useful seed-page result for this workflow: a system could preserve the page title and retrieved content as initial evidence, extract links from the returned page, and place only policy-approved candidates into its frontier. The proof does not establish browser rendering, automated link following, ranking, or tracing behavior for PagePith, so those capabilities should be validated separately for your integration.

Ship a small, inspectable first version

Start with one public domain, a maximum depth of two, a low page budget, read-only retrieval, and a required evidence record for every conclusion. Test it against pages with redirects, repeated links, irrelevant navigation, client-rendered content, and prompt-injection-like text. Review the trace before increasing autonomy.

A reliable link-following agent is not defined by how far it can browse. It is defined by whether it can explain what it followed, why it followed it, what it found, and why it stopped.

Ready to prototype a bounded retrieval step for your agent? Create a PagePith account.

Sources

  1. OpenAI Agents SDKOpenAI
  2. Playwright NavigationsMicrosoft Playwright
  3. Playwright BrowserContextMicrosoft Playwright
  4. Robots Exclusion Protocol, RFC 9309Internet Engineering Task Force
  5. OWASP Top 10 for LLM Applications 2025OWASP GenAI Security Project
  6. OpenAI Agents SDK TracingOpenAI
Build a Link-Following AI Research Agent · PagePith