Audra vs Screaming Frog: XML Sitemap Audit Guide
A practical XML sitemap audit comparison covering Audra, Screaming Frog, Sitebulb, Google Search Console, and lightweight sitemap checkers.
· 15 min read
A sitemap can be perfectly valid XML and still send search engines to redirected, noindexed, canonicalized, or 404 URLs. A useful XML sitemap audit therefore determines whether a sitemap contains the right URLs—not simply whether sitemap.xml loads—and gives agencies and site owners a repeatable way to document the fixes.
Screaming Frog’s XML sitemap tutorial is a strong starting point because it distinguishes sitemap-only URLs, crawl-only URLs, and non-indexable sitemap entries. But the best workflow goes further: validate the file, verify every listed URL with a crawl, and reconcile the findings with Google Search Console indexing evidence. This guide compares Audra, Screaming Frog SEO Spider, Sitebulb, Google Search Console, and lightweight sitemap checker tools.
| Tool or approach | Best at | Sitemap-specific depth | Pricing model as checked September 2026 | Ideal use case |
|---|---|---|---|---|
| Audra | Local full-site SEO, performance, accessibility, link, and AI visibility reporting | Use alongside a sitemap URL list and Search Console evidence | $19 one-time desktop purchase; no subscription | Client-ready, repeatable site audits on macOS or Windows |
| Screaming Frog SEO Spider | Detailed URL-level crawling and exports | High: sitemap, crawl, orphan, indexability, and duplication comparisons | Free up to 500 URLs; £199 per user/year for paid features | Technical SEO specialists and custom investigations |
| Sitebulb | Guided visual sitemap diagnostics and reporting | High: redirects, errors, noindex, orphan pages, Search Console connections | 14-day trial; plan pricing varies | Consultants who want prioritized explanations and visuals |
| Google Search Console | Google’s sitemap processing and page-indexing evidence | Medium: submitted sitemap status and indexing outcomes | Free | Confirming what Google has processed and indexed |
| Lightweight sitemap checker | Fast discovery and basic XML validation | Low | Usually free or usage-limited; varies by vendor | A quick first-pass sitemap test |
What an XML sitemap audit should prove
An XML sitemap is a crawling and discovery signal, not proof that its URLs will rank or even enter Google’s index. The Sitemap protocol itself says that it provides hints to crawlers and does not guarantee inclusion in search engines. (sitemaps.org) Google likewise states that it can take time to crawl sitemap URLs and may not crawl every submitted URL, depending on factors including site size, activity, and traffic. (support.google.com)
That creates three separate audit questions:
- Can search engines find and parse the sitemap? Check
robots.txt, the sitemap index, file syntax, status codes, host consistency, and protocol limits. - Are the listed URLs technically eligible? Each important sitemap URL should normally return an HTTP 200 status code, be indexable, and use the intended canonical tag.
- Does indexing evidence support the sitemap’s purpose? Compare submitted URLs with Search Console’s Page indexing report and investigate the important exceptions.
This is why a pass/fail sitemap checker is not enough for a client deliverable. It can catch a missing file or malformed XML, but it cannot reliably explain why 300 product URLs in a submitted sitemap resolve to canonicalized variants, why a migration left old redirects in the file, or whether Google is excluding key URLs for valid reasons.
For a broader framework around separating urgent technical blockers from lower-impact cleanup, see technical SEO audit prioritization. Sitemap defects are usually high priority when they affect important templates, sections, or many URLs—not merely because a checker labels them as warnings.
Audra vs Screaming Frog for an XML sitemap audit
Screaming Frog is the more specialized choice when the assignment is an intensive sitemap investigation. Its documented sitemap workflow can discover linked XML sitemaps through robots.txt or use a supplied sitemap destination, crawl the site, and then compare the sitemap’s URL set against crawl results. The paid version removes its 500-URL limit; Screaming Frog lists the annual licence at £199 per user as of September 2026. (screamingfrog.co.uk)
Audra is better understood as the practical reporting layer for a broader website review. It is a local desktop audit app for macOS and Windows that combines SEO, performance, accessibility, broken-link, and AI answer-engine visibility checks. Audra is sold as a $19 one-time purchase and runs audits and stores reports locally rather than requiring a subscription. (audra.greta.sh)
The practical difference
Use Screaming Frog when the task demands detailed sitemap-source analysis, custom crawl configurations, bulk exports, or granular investigation of a large URL list. It directly supports the sitemap-versus-crawl comparison that exposes URLs found only in a sitemap, URLs missing from it, non-indexable sitemap URLs, duplicate sitemap entries, and files over protocol limits.
Use Audra when a sitemap issue is one part of a wider client or owner audit. A redirecting URL in a sitemap rarely exists in isolation: it may also reveal poor internal linking, an old migration rule, slow destination pages, missing accessibility information, or weak on-page signals. Audra’s local, consolidated reporting is useful for turning that wider set of findings into an actionable deliverable without another monthly tool subscription.
The tools are complementary rather than mutually exclusive. An agency can use Screaming Frog for the forensic URL list, Google Search Console for Google-specific evidence, and Audra for the full-site audit and client-facing prioritization.
Find and validate XML sitemaps before crawling URLs
The first step is sitemap discovery. Start with https://example.com/robots.txt and look for one or more lines beginning with Sitemap:. Search Console also supports sitemap discovery through a submitted sitemap, while Google notes that a sitemap can be listed in robots.txt if the auditor does not have owner access to submit it in the Sitemaps report. (support.google.com)
Common paths such as /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml, and platform-generated variants are worth testing, but they should not be assumed to be canonical locations. A sitemap index may point to separate product, category, image, blog, or locale sitemap files. The index exists specifically to list multiple sitemap files and overcome the size limit of one file. (support.google.com)
A fast validation pass should check:
- The sitemap or sitemap index returns HTTP 200, without login, IP restrictions, or an accidental redirect chain.
- The XML uses the appropriate root structure:
<urlset>for a URL sitemap or<sitemapindex>for an index. - URL values are properly escaped and the file is UTF-8 encoded.
- Every listed URL belongs to the intended host and protocol, such as
https://www.example.com/rather than a mixture of HTTP, non-www, staging, or legacy domains. - Each sitemap contains no more than 50,000 URLs and is no larger than 50 MB uncompressed; use a sitemap index when either limit is exceeded. (sitemaps.org)
Lightweight sitemap checker tools are appropriate here. They provide a fast sitemap test for existence, discoverability, XML parsing, and initial URL extraction. They are not sufficient for the later URL-level decisions, because valid XML cannot reveal whether listed pages are canonical, useful, internally connected, or indexed.
Check every sitemap URL for status, canonical, and indexability
The central rule is simple: a standard XML sitemap should list the preferred URLs intended for indexing. Sitebulb summarizes the minimum condition as URLs that return a 200 response and are indexable. (sitebulb.com) In practice, an audit should add canonical consistency and business importance to that minimum.
For every URL extracted from the sitemap, inspect these fields:
| Check | Preferred result | Typical action when it fails |
|---|---|---|
| HTTP response | 200 OK | Remove 3xx, 4xx, 5xx, and timeout URLs; fix the underlying page where appropriate |
| Canonical tag | Self-referential or intended preferred URL | Replace duplicate variants in the sitemap with the canonical URL |
| Meta robots | No noindex on intended landing pages | Remove the URL from the sitemap or correct an accidental directive |
| robots.txt | Crawlable where crawling is required | Review whether the block is intentional; do not treat a blocked URL as a sitemap candidate |
| Final URL | HTTPS, preferred hostname, normalized parameters and trailing slash | Standardize generator rules and redirects |
| Internal discovery | Linked from the site where it should be | Add purposeful internal links or reconsider whether it belongs in the sitemap |
Broken and redirected sitemap URLs
A 404, 410, 5xx, timeout, or redirect in a sitemap is stale inventory. It wastes audit attention and gives search engines an unclear signal about which version matters. Do not simply replace a redirected URL with its destination automatically: first confirm that the final page is indexable, self-canonical, relevant, and the actual preferred URL.
For example, a sitemap could contain http://example.com/services/seo, which 301 redirects to https://www.example.com/seo-services/. The correct repair is normally to place the final HTTPS canonical URL in the sitemap, remove the old version, and check whether internal links still point to the legacy address.
Non-canonical and non-indexable sitemap URLs
A page may return 200 but still be unsuitable for the sitemap. Common examples include filtered faceted URLs, internal search results, paginated variants, printer pages, parameter URLs, duplicate color variants, pages with noindex, and pages canonicalized to a different URL. Screaming Frog’s tutorial specifically highlights non-indexable URLs in sitemaps as entries to remove or fix, not as harmless technical noise.
Compare sitemap URLs with the crawl to find gaps and orphans
The most valuable sitemap checks need two URL sets: URLs listed in the XML files and URLs discovered through a normal crawl. This is the feature that separates a proper XML sitemap audit from a standalone checker.
Screaming Frog’s sitemap analysis identifies:
- URLs in the sitemap: the URLs the site explicitly submits.
- URLs not in the sitemap: crawl-discovered, indexable pages that might be absent by mistake.
- Orphan URLs: URLs found only through the sitemap and not through internal crawling.
- URLs in multiple sitemaps: usually unnecessary duplication, even if it is not always harmful.
Sitebulb offers a similar comparison through its sitemap and crawl data, including sitemap-only URLs, redirects, errors, and indexation issues. (sitebulb.com)
An orphan URL deserves investigation, not automatic removal. A newly launched campaign landing page might be intentionally unlinked but still need indexing. More often, however, sitemap-only status reveals a weak information architecture: an important page has no crawlable internal route, so the sitemap is carrying discovery work that the site navigation should share.
Conversely, a high-value page discovered in a crawl but absent from the sitemap can indicate an incomplete generation rule. Check template patterns. If 150 new location pages are missing while 2,000 old location pages are present, the root cause is likely the CMS or generator configuration—not 150 individual URL edits.
For larger reviews, the better SEO audit plan guide helps place sitemap inventory checks alongside crawlability, internal links, performance, and page-level quality rather than treating them as an isolated task.
Reconcile the sitemap with Google Search Console indexing data
A crawler shows what the site serves now. Google Search Console adds evidence about what Google has processed and indexed. In the Sitemaps report, confirm submission history, sitemap processing status, and parsing errors. Google explains that submission means telling Google where the hosted file is; it does not mean uploading a file to Google. (support.google.com)
Then use the Page indexing report’s All submitted pages view. Google defines this as URLs listed in a sitemap or sitemap index submitted through the Sitemaps report or referenced in robots.txt. (support.google.com)
A practical reconciliation process is:
- Export or record the sitemap URL list and segment it by sitemap type, such as products, posts, locations, or images.
- Review Search Console sitemap errors and warning patterns first.
- Open Page indexing and filter to submitted pages; group exclusions by reason rather than reacting to isolated samples.
- Inspect representative high-value URLs with URL Inspection to review index status and Google-selected canonical.
- Compare the live crawl result with Google’s indexed-state evidence, then assign a fix owner: development, content, CMS, analytics, or SEO.
Search Console only gives examples for many Page indexing states, and the example list is capped at 1,000 URLs. It should therefore guide diagnosis rather than be treated as a complete sitemap export. (support.google.com) URL Inspection is the right tool for checking the indexed status and Google-selected canonical for a specific important URL. (support.google.com)
A concise Audra workflow for repeatable client audits
Audra works best when the sitemap review is folded into a complete website audit rather than presented as an isolated XML pass/fail test. The workflow below is deliberately tool-agnostic on sitemap extraction: use Screaming Frog, Sitebulb, or a basic sitemap checker to collect the sitemap URL list, then use the findings in the wider audit.
- Discover and validate the sitemap index. Record sitemap locations from
robots.txt, verify HTTP 200 responses, and note each child sitemap’s URL count. - Run the full-site audit in Audra. Review SEO, performance, accessibility, links, and AI answer-engine visibility on the same live domain.
- Match sitemap defects to broader site patterns. A non-canonical sitemap URL may correspond to sitewide parameter handling; 404 entries may signal a migration cleanup need; orphan pages may expose navigation gaps.
- Add Search Console evidence. Confirm whether the affected section is submitted, crawled, and indexed, and inspect representative commercial pages.
- Prioritize by impact. Lead with template-level issues, critical revenue or lead-generation pages, and widespread defects before cosmetic duplicate entries.
- Deliver a clear report. Separate immediate repairs—such as removing 404s and redirects—from structural work such as fixing sitemap generation rules or internal linking.
This local-first approach is particularly useful for agencies that need a report a client can understand, and for site owners who want a one-time tool instead of another recurring platform cost. It also leaves room to connect sitemap health with a broader brand discovery audit across search, AI, and social, since technically accessible pages still need clear content and entity signals to be discoverable.
Which XML sitemap auditing tool should you choose?
Choose a lightweight sitemap checker for a five-minute sitemap test: find the file, confirm that it loads, validate basic XML, and extract URLs. It is a triage tool, not an auditing system.
Choose Google Search Console whenever the question is “What has Google processed?” It is essential for sitemap submission status, parsing errors, submitted-page indexing patterns, and URL-level inspection. It does not replace a crawl because it does not provide a full live technical inventory.
Choose Screaming Frog SEO Spider for a deep one-off XML sitemap audit, migrations, custom extractions, and bulk spreadsheet analysis. The free version is useful for small sitemap lists up to 500 URLs; the paid licence is the practical choice for larger sites and full sitemap-plus-site-crawl comparisons. (screamingfrog.co.uk)
Choose Sitebulb when a consultant or agency needs guided issue explanations, visual prioritization, sitemap-versus-crawl checks, Search Console connections, and polished reporting. It is especially suitable when the audit recipient benefits from more context than raw crawl exports. (sitebulb.com)
Choose Audra when the sitemap is one component of a repeatable, local, client-ready website audit that also needs SEO, speed, accessibility, links, and AI visibility findings. For highly specialized sitemap parsing and source-level exports, pair Audra with Screaming Frog or Sitebulb rather than pretending that a general audit report should replace forensic crawler analysis.
Verdict
The strongest XML sitemap audit is a workflow, not a single tool. Start with discovery and syntax validation, crawl every sitemap URL for 200 status, canonical tags, and meta robots directives, compare sitemap and crawl inventories, then reconcile priority URLs with Google Search Console.
Screaming Frog remains the strongest fit for detailed technical sitemap investigations. Sitebulb adds guided diagnostics and visual reporting. Search Console supplies Google’s indexing evidence. Audra is the practical choice for teams that need to turn sitemap findings into a wider local-first audit and a report clients or site owners can act on.
FAQ
How do you audit an XML sitemap for SEO?
Audit the file in three layers: first verify discovery through robots.txt, XML structure, HTTP 200 access, and sitemap-index limits; then crawl each listed URL for response code, canonical tag, meta robots, and crawlability; finally compare the URL inventory with Google Search Console submitted-page indexing data. A valid file alone does not prove that its URLs are eligible or indexed.
How can you find an XML sitemap on a website?
Check the site’s robots.txt file first, because it may list one or more Sitemap: directives. Then test common sitemap paths and review Google Search Console’s Sitemaps report if access is available. Follow sitemap index files to their child sitemaps; a large ecommerce site may split products, categories, posts, and images into separate XML files.
What should be checked in an XML sitemap?
Check that the sitemap returns HTTP 200, uses valid XML and UTF-8 encoding, stays within 50,000 URLs and 50 MB per sitemap, and lists only intended preferred URLs. At URL level, review redirects, errors, canonical tags, meta robots directives, robots.txt blocking, hostname consistency, duplicate entries, stale content, and internal-link discovery.
How do you identify broken, redirected, non-canonical, or non-indexable URLs in a sitemap?
Export or crawl the sitemap URL list in Screaming Frog, Sitebulb, or another crawler. Filter by HTTP 3xx, 4xx, 5xx, and timeout responses; compare each URL’s canonical target; and review meta robots and robots directives. Sitemap entries should normally resolve directly to 200, indexable, preferred canonical URLs rather than to alternates or obsolete addresses.
How do you compare submitted sitemap URLs with Google Search Console indexing data?
Open Search Console’s Sitemaps report to confirm the submitted file and any processing errors. Then review Page indexing using the submitted-pages view, which focuses on URLs from submitted or robots.txt-discovered sitemaps. Use URL Inspection for critical examples to check indexed status and Google-selected canonical. Treat the report’s examples as diagnostic samples, not a complete list.
Which XML sitemap auditing tool is best for a one-off or recurring audit?
For a one-off technical investigation, Screaming Frog offers the most granular sitemap-versus-crawl analysis. Sitebulb is strong for guided recurring audits and visual reports. Google Search Console is necessary for Google-specific processing and indexing evidence. Audra suits recurring local-first site audits where sitemap issues need to be prioritized alongside performance, accessibility, links, SEO, and AI visibility findings.
Sources
- https://www.screamingfrog.co.uk/seo-spider/tutorials/how-to-audit-xml-sitemaps/
- https://www.screamingfrog.co.uk/seo-spider/pricing/
- https://support.google.com/webmasters/answer/7451001?hl=en
- https://support.google.com/webmasters/answer/7440203
- https://support.google.com/webmasters/answer/9012289?Hl=en
- https://www.sitemaps.org/protocol.html
- https://sitebulb.com/product/xml-sitemaps/
- https://audra.greta.sh/
- https://www.sitemaps.org/
- https://support.google.com/webmasters/answer/12818558?hl=en