← Back
ai searchai crawlerstechnical seojavascript seowebsite audits

AI Crawler Content Visibility: What 300 Sites Actually Serve

A reproducible audit method shows what a website serves before rendering, where page meaning is lost, and how agencies can prioritize fixes with evidence.

· 13 min read

A 300-site experiment reported a median of 2,592 extracted tokens per page, with roughly half not attributable to main body content. An AI crawler content visibility audit turns that finding into a concrete decision: it identifies whether a priority URL passes, needs review, or fails because the server response does not contain enough page-specific meaning.

The useful comparison is not “what does the browser show?” alone. It is what a controlled request receives over the wire—status code, redirects, headers, HTML, robots instructions, and any negotiated response—compared with the meaningful content a browser presents after rendering. That evidence helps site owners fix delivery problems before assuming AI-search visibility is solely an editorial problem.

Set an operational AI crawler content visibility score

A useful audit needs a decision rule rather than a general claim that a page is “AI-ready.” For each priority URL, assign a score out of 100 from four recorded checks:

  1. Access: 30 points. The canonical URL returns a successful final response without an unintended login wall, WAF challenge, timeout, or loop.
  2. Source-content availability: 30 points. The initial response contains the page’s primary answer, H1, and enough page-specific text to establish its purpose.
  3. Meaning clarity: 25 points. Main content is distinguishable from navigation, consent controls, repeated calls to action, and footer text.
  4. Variant integrity: 15 points. Any HTML, Markdown, mobile, locale, or bot-targeted variant preserves material facts and is safely cached.

A score of 85-100 is a pass: no material delivery issue was observed under the recorded test conditions. 60-84 needs review because access works but rendering dependence, excessive boilerplate, or a variant difference could affect interpretation. Below 60 is a fail for that URL because meaningful content is missing, blocked, or materially altered.

This is an audit rubric, not a prediction of whether a named answer engine will cite a page. The 300-site study did not measure citations, rankings, or all crawler implementations. It measured delivered content, which is a narrower and reproducible technical question.

What the 300-site experiment actually found

The TechSEO community post and its accompanying results documentation examined 300 public sites with an open-source content-parity approach. It used the cl100k_base tokenizer and reported a median of 2,592 tokens in extracted page text. The study also reported that approximately half of extracted material was not main body content and that headers, navigation, and footers represented about one fifth of the median response.

Those are sample results, not web-wide prevalence figures. They do not show that half of all websites are hidden from LLMs, that every AI crawler ignores JavaScript, or that every non-body token is wasted. A global navigation menu can be useful to users and to some machine tasks.

They do establish a practical concern: a page that looks complete in Chrome may still send an automated fetcher a large quantity of template material relative to the explanation, evidence, product details, or answers that make the URL distinct. The study also identified seven Markdown variants, with a reported median of 962 tokens, making response variants a specific area worth testing rather than assuming.

For agencies, the result is best used as a benchmark for investigation. It supports asking whether a client’s important pages deliver identifiable, page-specific content in the source response—not treating the study’s token counts as universal targets.

Define the three versions of a page

Browser-rendered content

A human browser can load CSS, execute JavaScript, complete client-side hydration, make API requests, and display a consent-selected state. A product category may show 48 products after an application calls /api/catalogue, even if the original HTML only contains <main id="app"></main>.

That browser view matters for user experience, but it is not proof that every AI search bot, content licensing crawler, or social-preview bot receives the same page. Different systems can use different rendering, caching, request-header, and anti-bot handling.

Initial server response

The initial response is the most defensible baseline. It includes the final HTTP status, redirect chain, headers, and body supplied to a request before client-side JavaScript changes the document.

For an article, a strong initial response usually includes the title, H1, introductory answer, core headings, and principal text. It does not need to duplicate every interactive enhancement. The important question is whether the source response establishes what the page is about and includes the information needed for the page’s central purpose.

Negotiated or classified variant

A CDN, origin, or application can alter content based on the Accept header, user agent, cookie state, geography, device, or bot classification. Markdown, AMP, locale-specific, and preview-oriented variants can all be useful, but they need separate evidence.

The audit objective is parity of meaning, not byte-for-byte identity. A Markdown response can properly remove decorative navigation, animations, and layout wrappers. It should not silently omit a pricing table, product specification, author credential, limitation, citation, or conclusion that changes the page’s meaning.

The experiment directly supports investigation into template-heavy responses and Markdown variants. The following are recommended diagnostic categories based on standard web-delivery failure modes; they were not quantified as prevalence findings across the 300-site sample.

  • JavaScript-only primary content: A raw response has a loading shell while a rendered browser contains a 900-word guide, FAQ, or product table. This is a source-content availability failure because access depends on rendering.
  • Consent and interstitial layers: A cookie interface, age gate, newsletter prompt, or region selector appears before the main content. It may not block a crawler, but it can dilute extracted text and needs a clean-cookie test.
  • Blocked assets or APIs: HTML may return 200 OK while the script, API, or image endpoint carrying essential meaning returns 403, 429, 503, or a timeout.
  • Redirect and canonical mismatch: A request traverses http to https, non-www to www, locale routing, or a security redirect before landing on a body different from the declared canonical URL.
  • Bot-control response differences: The same URL may be allowed in robots.txt but challenged by Cloudflare, a WAF, or origin access rules.

A report should label each result accurately: “observed on this tested URL,” “reported by the 300-site study,” or “recommended check.” That distinction avoids presenting general technical SEO advice as if it were a measured population statistic.

Robots.txt, headers, and server blocks are different layers

Websites blocking AI crawlers may be making a deliberate commercial or editorial choice. Others may unintentionally apply a broad User-agent: * directive, inherit a hosting template, or allow a robots rule while a CDN blocks the actual request. These are different technical states and should not be collapsed into one red “blocked” label.

First, inspect robots.txt for the exact tested user agent and any broad wildcard rule. Google describes robots.txt as a protocol used by compliant crawlers to determine which URLs they may crawl; it is not an access-control mechanism that prevents a direct HTTP request.

Second, request the URL and preserve the final response. A URL can be allowed by robots rules but return a 403 Forbidden, a 429 Too Many Requests, a Cloudflare challenge, or a login screen. Conversely, a disallowed URL may return normal HTML to a direct fetch.

Third, record X-Robots-Tag headers and HTML robots meta tags separately. Google documents directives such as noindex, nosnippet, and max-snippet for Google’s systems. Their presence should be reported precisely; an audit should not assume every AI system interprets them in the same way.

Finally, distinguish claimed from verified bot traffic. A user-agent string can be spoofed. Log reviews should use vendor-supported verification methods where available, then state what was verified rather than assuming every request labelled as an AI bot was genuine.

Use one reproducible methodology

The strongest audit does not impersonate every proprietary crawler or claim certainty about unseen retrieval systems. It records a small, repeatable request matrix that a developer can rerun after a fix.

Start with 20-50 URLs: the homepage, key service pages, highest-value product or category templates, important articles, location pages, and URLs associated with AI-search goals. Include at least one example of every CMS or template family.

For each URL, capture:

  1. Request timestamp, URL, request headers, user agent, cookie state, and test region when relevant.
  2. Redirect hops, final status, content type, response size, cache headers, Vary, and CDN/WAF indicators.
  3. Applicable robots.txt directive, page-level robots directives, login requirement, rate limit, or challenge response.
  4. Extracted raw text, rendered-browser main text, H1 and major headings, canonical URL, internal links, and structured-data presence.
  5. Any alternate representation requested through a documented header, such as Accept: text/markdown.

The strongest worked example is simple: a commercial page returns 112 words in raw HTML and no pricing criteria, but its rendered state contains a 680-word comparison and pricing table. That URL fails source-content availability until the essential comparison copy and data are server-rendered, statically generated, or otherwise included in the initial response.

This workflow is a recommended practice. Audra’s established role is to combine AI answer-engine visibility checks with technical SEO, performance, accessibility, and link audits in a local desktop application. The raw-versus-rendered comparison, response capture, and parity review described here should be treated as an audit workflow teams can reproduce, not as a claim about a specific Audra feature.

Test JavaScript rendering without assumptions

Googlebot provides a useful comparison but not a universal model. Google’s JavaScript SEO documentation explains that Google can process JavaScript and also warns that blocked resources can prevent effective rendering. That supports testing dependency access; it does not establish how an LLM crawler, Meta crawler, or another AI search bot renders a particular framework.

Use two controlled retrieval modes for every high-value page:

  • Raw fetch: capture the HTML response before JavaScript execution.
  • Rendered browser: load the same canonical URL in a controlled browser session, wait for the relevant page content, and extract the visible main content.

The decision rule is based on material meaning. If the rendered page adds only an accordion animation or a product-image carousel, the raw response may still pass. If it adds the primary answer, product specifications, article body, or key internal links, the URL needs review or fails depending on the scale of the omission.

Meta crawlers, Googlebot, AI training crawlers, answer-engine retrieval agents, and social preview fetchers should not be treated as one bot class. Their declared purposes and practical request patterns vary, which is why a site-specific response record is more useful than a blanket assertion about “AI bots and JavaScript.”

Markdown variants require cache safety and content parity

The study’s seven Markdown variants are the clearest technical example of why alternative machine-readable output must be tested. It reported that six of seven implementations lacked Vary: Accept, and that three of seven dropped actual body content, with those omissions associated with one platform.

Vary: Accept tells shared caches that the representation can vary according to the request’s Accept header. Without it, a cache may not reliably separate HTML and Markdown responses. The appropriate report finding is not “Markdown is bad” but “this URL serves variants without sufficient cache variation” when that is observed.

The prior version of this article said Cloudflare’s current Markdown for Agents documentation showed the correct Vary: Accept pattern. That statement should not be published as written. The supplied study specifically identified problems on Cloudflare’s own documentation page at the time it was tested, including the cache-variation issue. Documentation and product behavior can change, and no current Cloudflare documentation claim should override a recorded test result without a dated, independent recheck.

For each tested variant, compare final status, Content-Type, Vary, cache headers, title, H1, main answer, tables, prices, citations, limitations, FAQs, and internal links. A lower token count is beneficial only when it removes interface noise without removing decision-relevant facts.

Prioritize fixes by lost meaning and business intent

Cloudflare controls, Fastly configurations, AI crawler traffic, mixed-use crawlers, and pay-per-crawl models are relevant policy topics. They do not change the immediate technical decision: does the business intend a given agent category to access this URL, and does the response provide an accurate representation?

Use three report priorities:

  • Critical: a high-value URL returns a challenge, 403, timeout, empty application shell, redirect error, login requirement, or materially incomplete alternate representation.
  • High: main text exists but essential data depends on an unreliable API, template material dominates extraction, or HTML and Markdown differ in facts or cache behavior.
  • Improve: semantic landmarks are weak, repeated navigation is excessive, or evidence, author details, and FAQ headings are difficult to isolate.

A publisher may intentionally restrict a training-oriented crawler while permitting search-focused crawling, or may choose a paid-access model. That is a valid policy outcome, not an audit error. The report should identify whether the implementation matches the stated policy.

When access is sound, content quality remains a separate question. A practical evidence scorecard for brands winning AI search can help assess whether accessible pages also supply specific, supported, citable information.

Build a client report that developers can act on

Each issue should contain an artifact, not merely a score. Show the tested URL, timestamp, response status, redirect chain, applicable robots directive, key headers, raw-text excerpt, browser comparison, and a named technical cause.

A concise issue card might read:

> URL: /pricing/enterprise > Observed result: Raw HTML contains 112 words and no pricing criteria; the rendered page contains a 680-word comparison and pricing table. > Cause: Client-side API rendering with no source HTML fallback. > Priority: Critical for a primary conversion page. > Fix and validation: Include comparison copy and pricing data in initial HTML, then rerun the identical raw and rendered tests.

For a bot-control issue, include the exact robots rule and observed response separately. For a Markdown issue, include both bodies, token counts, Content-Type, Vary, and cache headers. This makes ownership clear across SEO, development, CDN, and content teams.

The approach fits a local-first audit process because teams can retain response evidence and client reports without treating a subscription dashboard as the sole source of truth. It also complements a broader SEO audit plan that prioritizes fixes: technical evidence should determine what gets fixed first.

FAQ

What is an AI crawler?

An AI crawler is an automated web client that fetches content for purposes such as answer retrieval, search, training, licensing, previews, or agent workflows. It is not a single standardized technology. Googlebot, Meta-related crawlers, and other AI search bots can differ in user agents, rendering behavior, request headers, verification methods, and policy compliance.

What content do AI crawlers actually receive from a website?

They receive the response supplied under their request conditions: status code, redirect chain, headers, and body. That may be initial HTML, a JavaScript-rendered result, a Markdown variant, a preview representation, or a challenge page. Browser appearance is useful evidence, but it does not prove what a non-browser fetcher received.

How much website content is hidden from AI crawlers because of JavaScript or bot controls?

There is no defensible universal percentage. The 300-site experiment reported substantial non-body extracted text in its specific sample, but it was not a census of all websites or all crawlers. Site owners should test representative URLs for content that appears only after JavaScript or is blocked by robots rules, WAFs, CDNs, login requirements, or failed APIs.

Can robots.txt block AI crawlers, and do sites unintentionally block them?

Robots.txt can instruct compliant crawlers not to fetch specified paths, including through broad User-agent: * rules or named-agent directives. Unintended restrictions can result from copied templates and broad wildcards. Robots rules are separate from server enforcement, so an audit must also request the URL and record whether the server returns content, a denial, or a challenge.

How can a site owner test what AI crawlers can see without relying on assumptions?

Use a documented request matrix on 20-50 priority URLs. Capture raw HTML, status codes, redirects, headers, robots directives, and any alternate representation; then compare the results with a controlled rendered-browser view. Save the evidence and rerun the same checks after deployment. This establishes what the site served under recorded conditions without claiming to replicate proprietary systems perfectly.

Sources