← Back
orphan pagestechnical seointernal linkingsite auditseo tools

Find Orphan Pages: Audra vs Screaming Frog Audit Workflow

A dependable orphan-page audit compares the internally linked crawl with sitemap, analytics, Search Console, and, where available, server-log URL inventories before recommending a fix.

· 15 min read

A hypothetical 500-page site can produce three conflicting URL lists: 420 URLs in a crawl, 470 in XML sitemaps, and 450 landing-page URLs in analytics. The gap is where teams can find orphan pages that a crawl alone cannot reach—and turn a noisy export into a defensible list of pages to link, consolidate, redirect, noindex, or remove.

Screaming Frog’s documented orphan-page workflow combines a standard crawl with XML sitemaps, Google Analytics, and Google Search Console. Audra is a local-first desktop audit agent that adds technical SEO, links, performance, accessibility, best-practices, and AI answer-engine visibility checks to the wider audit and reporting process. The practical choice depends on whether the immediate need is deep URL reconciliation or a broader client-facing site audit.

MethodNative discovery strengthWhat requires exports, connections, or another sourcePricing contextBest fit
AudraLocal audits of technical SEO, links, performance, accessibility, best practices, and AI visibilityA separate known-URL inventory is still needed to identify URLs absent from the crawlNo subscription is described in Audra’s product brief; current purchase price is not stated hereBroader local audits and client reporting
Screaming Frog SEO SpiderCrawled internal-link structure; documented sitemap, Analytics, and Search Console orphan workflowRelevant APIs and source data must be connected; a licence is required for the full workflowLicence requirement documented; current price and crawl limits are not stated hereDetailed orphan-page investigation
Semrush Site AuditSite-audit crawl and orphan-page-oriented guidanceAccount configuration and connected data scope varySubscription details vary and are not compared hereExisting Semrush users
Free orphan page checkerFast initial screen of a limited public site viewComplete sitemap, analytics, Search Console, CMS, and log reconciliationVaries by providerSmall-site triage
Server-log analysisRequested URL evidence, including visits from verified GooglebotLog access, parsing, and bot verificationHosting and analysis arrangements varyLarge or complex sites

What an orphan page means—and what it does not

An orphan page is a URL with no meaningful inbound internal link in the site structure observed by the audit. Put more simply: the page may exist, but the crawler cannot find an internal path to it from the selected crawl starting point.

That is deliberately a crawl-based definition, not a claim that nobody can reach the URL. A page can still appear in an XML sitemap, receive visits from an email campaign, earn search impressions, or attract external backlinks. The Google Search Central guidance on crawlable links explains why normal HTML links matter for discovery, but an orphan finding still needs validation before it becomes a remediation task.

A useful audit separates four outcomes:

  • Likely true orphan: the URL appears in a sitemap, analytics, Search Console, CMS export, or logs but has no verified internal route.
  • Weakly connected page: a link exists, but it is buried or poorly contextual. This is an internal-linking quality problem, not necessarily an orphan.
  • Intentional exception: a thank-you page or short-lived campaign destination may be intentionally absent from the main architecture.
  • Candidate needing validation: the audit did not observe a path, but the team has not yet checked templates, navigation states, or crawl configuration.

The distinction matters because “zero inlinks” is an observation, while “this page should be linked” is a strategic decision.

Audra vs Screaming Frog for an orphan-page audit

Screaming Frog is the more direct specialist when the core deliverable is an orphan-page candidate list. Its documented tutorial sets out a process that brings together a crawl, XML sitemaps, Google Analytics, and Google Search Console. Those sources expose URLs that a standard link-following crawl may never encounter.

Audra serves a broader role. It is a local desktop AEO audit agent for macOS and Windows that assesses AI answer-engine visibility alongside technical SEO, performance, accessibility, best practices, and links. Its product brief also describes client-ready reports without a subscription. For an agency, that makes Audra useful for explaining the surrounding quality and implementation context after orphan candidates have been identified.

What each option can establish

Neither product should be represented as proving every URL on a site is orphaned from crawl data alone. The evidence differs:

  • Screaming Frog: can reconcile the crawl against connected sitemap, Analytics, and Search Console sources through its documented workflow.
  • Audra: can audit the crawlable site experience and related link and technical conditions, but the brief does not claim a native import-and-reconcile orphan-page feature.
  • Semrush Site Audit: is relevant to teams already using its site-audit environment; its public orphan-page guidance should be checked against the account’s configured crawl and connected data.
  • Free checkers: can flag isolation quickly, but should not be treated as a full URL inventory without first-party data.
  • Logs: can add evidence that a URL was requested, but a request alone does not establish that the URL is valuable or should be indexed.

This is why a practical agency workflow can use Screaming Frog for source reconciliation and Audra for the broader local audit and client report. The tools solve adjacent parts of the job rather than identical ones.

Why a crawl alone cannot find orphan pages

A normal crawler follows URLs exposed through links. If a discontinued product page is no longer linked from category pages, a homepage-led crawl may not discover it at all. Yet that page could still return a 200 response, remain in a sitemap, receive visits from external links, or show up as a landing page in analytics.

For example, consider /guides/old-widget-installation/:

  1. It was removed from the help-centre navigation during a redesign.
  2. It remains in the XML sitemap.
  3. It earned organic visits during the selected reporting period.
  4. No URL in the internal crawl links to it.

The crawl identifies the linked site graph; the sitemap and traffic sources identify a wider URL universe. The candidate becomes more credible only after someone checks whether a meaningful internal link exists but was missed by the crawl setup or is absent from the audited section.

This mirrors BrightEdge’s broad approach in its orphan-page guidance: compile a complete URL inventory, then crawl the site to identify the URLs with no inbound internal links. It is a comparison exercise, not a single-filter exercise.

How to find orphan pages with source reconciliation

The following sequence retains the useful part of the Screaming Frog workflow while avoiding the false promise that any one data source is complete.

1. Crawl the internally linked site

Start from the canonical homepage or a defined site section. Export the crawl URL list and capture practical fields such as final URL, HTTP status, crawl depth, internal inlink count, referring URL, and anchor text where the tool provides them.

This becomes the linked crawl set. It answers one focused question: which URLs did the configured crawler reach through the observed internal site structure?

Before comparing sets, document the crawl scope. A subfolder-only crawl, a restricted user agent, or an incomplete rendering configuration can change the result. Those are audit conditions, not proof that a page is disconnected.

2. Export XML sitemap URLs

Gather URLs from the active sitemap index and its child sitemaps. Google’s sitemap documentation is useful context: sitemaps are a way to tell search engines about site URLs, not a guarantee that every listed URL is current or indexable.

Normalize the lists before comparison. At minimum, agree how the audit handles:

  • http versus https
  • www versus non-www
  • trailing slashes
  • redirected URLs
  • tracking parameters and URL encoding

A sitemap-only URL can be a valuable forgotten page, a redirect that was never removed from the sitemap, or a CMS-generated URL that should not be there. The source identifies a candidate; it does not dictate the fix.

3. Add analytics and Search Console page lists

Analytics landing-page data identifies URLs that users reached in the chosen date range. Search Console performance data identifies pages associated with Google Search impressions or clicks in the selected report. Screaming Frog’s tutorial specifically includes both sources in its orphan-page process.

Choose a time window that fits the business. A 30-day window may suit an active retailer, while a seasonal service business may need a longer period or a comparable season. There is no universal reporting range supported by the sources because traffic patterns differ.

The basic comparison is:

URL in sitemap, analytics, or Search Console minus URL in linked crawl = orphan-page candidate

Google Search Console should not be treated as a complete orphan-page checker. It is a valuable source of search-related evidence, but absence from a performance export does not prove a URL does not exist. Google’s Search Analytics data guidance also explains that exports have practical data limits and aggregation considerations.

4. Add logs when they are available

Server logs are particularly useful on large publishers, marketplaces, and sites with extensive historical URL patterns. They can show that a URL was requested by a crawler even when the current internal crawl does not include it. Botify’s orphan-page discussion makes the same useful distinction: crawl data shows linked structure, while logs can expose URLs visited by Googlebot outside that structure.

Use verified bot data rather than assuming that any request with a Googlebot-looking user agent is genuine. Google’s Googlebot documentation provides its own verification guidance. Where reliable logs are unavailable, say so in the report rather than implying that crawler activity was assessed.

5. Validate and classify each candidate

Check the candidate URL manually before recommending a change. Confirm its status code, intended audience, content quality, current sitemap inclusion, and whether a relevant parent or topic page exists. Then classify the action rather than creating a blanket “add links” task.

How to fix orphan pages after they are found

The right fix depends on purpose and value. A useful decision model has five outcomes.

Add contextual links when the page is current, useful, and aligned with an existing topic or product journey. A setup guide, for example, may belong on the relevant product page, support hub, and troubleshooting article. The goal is not merely to increase an inlink count; it is to give users an understandable route.

This is also where internal-link decisions affect broader content architecture. Pages with overlapping subjects need distinct roles and deliberate connections. The guide to when keyword cannibalization can improve rankings is relevant when two similar pages serve genuinely different intents rather than being accidental duplicates.

Consolidate or redirect pages with no independent role

If an orphan duplicates a stronger page, merge the useful material into the stronger destination and use an appropriate permanent redirect where the old URL has a clear successor. Update internal references and remove retired URLs from the sitemap.

A replaced product may also merit a redirect to its successor or closest relevant category. A generic redirect to the homepage does not solve the user’s original task.

Noindex or remove intentional exceptions

Some isolated URLs are legitimate: campaign thank-you pages, login routes, or expired paid-media destinations may not need an internal discovery path. The action might be noindexing, restricting access, returning an appropriate status when the page is retired, or leaving it available for its intended channel.

The report should state the reason. “No internal links” alone is not enough evidence to remove a live page.

Repair the publishing process

Repeated orphans often indicate a workflow problem. Add publication checks such as: is the page assigned to a hub, is it linked from an appropriate parent page, and is its sitemap inclusion intentional? This matters for trust-oriented pages too. An about page SEO strategy is more credible when the company history and proof pages connect to relevant services, people, and supporting content instead of sitting as isolated brochure content.

Tool comparison: discovery versus audit context

The compact matrix below separates documented discovery mechanisms from work that remains manual or dependent on other systems.

Tool or approachCan natively crawl internal links?Can use sitemap, Analytics, and Search Console URL sources?Can add server-request evidence?Appropriate conclusion
Screaming Frog SEO SpiderYesYes, in its documented orphan-page workflowNot established by the cited orphan-page tutorialStrong candidate-list workflow when sources are connected
AudraAudits site links and technical conditionsNot specified in the supplied product briefNot specified in the supplied product briefBroader local audit and reporting context; reconcile external lists separately
Semrush Site AuditYes, as a site-audit crawlerOrphan-page handling depends on the configured product workflowNot established hereUseful for existing platform users, subject to configuration
Free checkerUsually limited public checkingGenerally requires outside data for complete reconciliationNoInitial screen only
Log analysisNo, it is not a link crawlRequires comparison with a crawl and URL inventoryYes, when logs are available and bot traffic is verifiedSupplemental evidence for complex sites

Sitechecker and similar free tools can help a site owner begin an investigation, while Semrush can be convenient for teams already operating audits there. Neither category removes the need to validate the candidate URL against first-party inventories. Feature availability, plan limits, and integrations should be checked in the vendor’s current documentation rather than assumed from a generic comparison.

Which should you choose?

Choose the method according to the deliverable and the evidence available.

  • Choose Screaming Frog when the central task is to identify orphan-page candidates by comparing an internal crawl with XML sitemaps, Google Analytics, and Google Search Console data.
  • Choose Audra when the team also needs a local desktop audit across AI visibility, technical SEO, performance, accessibility, best practices, and links, with client-ready reporting and no subscription described in the product brief. Pair it with exported URL lists for orphan-page reconciliation.
  • Choose Semrush Site Audit when the organisation already uses Semrush and has verified that its configured audit and integrations meet the needed workflow.
  • Choose a free checker for a quick small-site signal, then validate every finding with crawl and first-party data.
  • Choose log analysis when server access exists and Googlebot’s observed requests materially affect the decision, especially on large or legacy-heavy sites.

For many agencies, the practical sequence is diagnosis with a reconciliation-capable crawler, broader quality assessment and reporting in Audra, implementation by the site team, and a follow-up crawl to confirm that the selected pages now have the intended internal route.

Verdict

The reliable way to find orphan pages is to compare the linked crawl against other URL inventories, not to treat one crawler report as complete. Screaming Frog is the more direct choice for the documented sitemap, Analytics, and Search Console reconciliation workflow. Audra is a complementary local-first audit option when the orphan-page decision must sit alongside technical SEO, link, performance, accessibility, best-practice, and AI-visibility findings.

The strongest report identifies the source of each candidate, records the validation result, and recommends one clear action: link, consolidate, redirect, noindex, remove, or leave intentionally isolated.

FAQ

How can I find orphan pages in Screaming Frog?

Run a standard crawl, then follow Screaming Frog’s documented process for adding XML sitemap URLs and connecting Google Analytics and Google Search Console data. Compare URLs found in those sources with the internally linked crawl set. The output is a candidate list, so inspect each URL’s inlinks, status, page purpose, and any relevant crawl-scope limitations before making changes.

What does an orphan page mean?

An orphan page is a URL with no meaningful internal link path in the site structure observed by the audit. It can still be indexed, listed in an XML sitemap, receive analytics traffic, or have external backlinks. The term describes disconnection from the audited internal-link graph; it does not automatically mean the page is broken or should be deleted.

Is there a way to see all the pages of a website?

No single source reliably shows every URL. A crawl shows linked URLs, XML sitemaps show declared URLs, analytics shows visited landing pages, Search Console shows search-performance evidence, and server logs can show requested URLs. A CMS export can add another inventory. Reconciling these sets is the closest practical way to map a site’s full URL universe.

Can Google Search Console find all orphan pages?

No. Search Console can contribute pages with search impressions or clicks and useful indexing or sitemap context, but it does not provide a complete orphan-page report. A URL absent from its performance data may still exist or simply have no recorded activity in the chosen period. Compare Search Console exports with crawl and sitemap inventories instead.

How should orphan pages be fixed after they are found?

First decide whether the URL has a useful, distinct purpose. Add contextual internal links to valuable pages, consolidate duplicates into a stronger destination, redirect replaced pages to a relevant successor, and noindex or remove pages that are intentionally unsuitable for search. Finally, correct sitemap entries and publishing processes so the same isolation problem does not recur.

Sources