Raw HTML vs Rendered DOM: Website Visibility to LLMs
Website visibility to LLMs depends on more than whether a browser can display a page: teams need to compare raw HTML, rendered content, crawl access, structured data and real AI-answer exposure.
· 16 min read
A product page can look complete in Chrome yet return an almost empty document before JavaScript runs. That gap is the practical problem behind website visibility to LLMs: a site owner needs evidence of what answer engines and their crawlers can fetch, understand and potentially use—not just reassurance that the page works for human visitors.
Sitebulb’s August 2025 guide, The Invisible Web, usefully highlights the raw-HTML versus rendered-HTML problem. Its central warning is worth keeping, but it needs one qualification: fetching a page, rendering it, indexing it, training on it and citing it in an AI answer are separate stages. No single test can prove every LLM will know, quote or recommend a page.
| Dimension | Raw HTML inspection | Rendered DOM inspection | Audra visibility audit |
|---|---|---|---|
| Primary question | What arrives in the first server response? | What appears after the browser runs JavaScript? | Which technical and content layers could limit AI-answer visibility? |
| JavaScript content | Usually absent if injected client-side | Usually present when scripts run successfully | Flags the underlying SEO, performance, accessibility and link conditions alongside AI visibility checks |
| Structured data | Can confirm JSON-LD or microdata in source | Can reveal markup added after rendering | Checks page-level evidence in a wider audit workflow |
| Crawl directives | Can inspect robots.txt and meta robots | Does not prove a bot is permitted or willing to crawl | Helps teams review access and technical signals together |
| Best use case | Finding content LLM crawlers may miss | Diagnosing browser and Google-style rendering gaps | Producing a local-first, client-ready audit without a subscription |
| Main limitation | Does not show user-facing rendered experience | Does not show what every AI crawler receives | Cannot guarantee a future AI answer, citation or model-training outcome |
Raw HTML vs rendered DOM: the visibility split
The most useful starting point is to treat a URL as two different documents:
- Response HTML is the initial source returned by the server. In Chrome, this is close to what appears in “View Page Source.”
- Rendered DOM is what exists after the browser downloads scripts, executes JavaScript, makes API calls and modifies the page.
For a server-rendered product page, the initial response may already include an H1, price, specification table, internal links and FAQ text. For a client-side React application, the response might contain a root <div>, JavaScript bundles and little meaningful copy until the app runs.
Google documents that its systems can process JavaScript, while also recommending that sites avoid relying on JavaScript for critical links and content where possible. That makes rendered inspection relevant for Google Search and Google-connected experiences. It does not establish that every LLM crawler, answer engine or training pipeline renders pages in the same way, at the same depth, or at the same time.
A reliable audit therefore compares both versions. If the raw HTML includes a page title, canonical URL, H1, core proposition and crawlable internal links, the site has a stronger baseline. If those essentials appear only after client-side rendering, the page has a dependency that should be tested rather than assumed safe.
Sitebulb calls attention to three especially consequential examples: an H1 added only after rendering, JavaScript-modified headings and critical body copy loaded after scripts run. For an ecommerce category page, missing raw HTML could mean that the initial document has neither “Women’s Trail Running Shoes” nor links to the actual products. A human sees the content; a simple fetcher may not.
What LLMs see on a website—and what they can miss
It is tempting to frame this as “LLMs only read raw HTML.” That is too absolute. Different providers use different crawling, browser, retrieval and partnership systems, and those systems change. A more accurate model is that raw HTML is the lowest-friction representation of a page, while rendered content has additional technical conditions.
Content that is comparatively durable
The following items are generally easier to inspect and fetch when they are present in the first HTML response:
- A descriptive
<title>, H1 and introductory copy. - Body text explaining products, services, locations, authorship or policies.
- Standard
<a href>internal links to important pages. - JSON-LD structured data included in the document source.
- Meaningful image
alttext and text alternatives for non-text content. - A visible HTML table for prices, specifications, eligibility criteria or comparisons.
These elements do not force an AI answer engine to cite the site. They simply reduce the risk that the basic facts are absent from a first-pass fetch.
Content that frequently needs closer testing
The risk rises when content depends on browser events, user state or late API responses. Examples include:
- Tabs that load product details only when clicked.
- “Load more” results whose URLs are not exposed in ordinary links.
- Chat widgets containing the only service explanation or lead qualification details.
- Customer reviews injected from a third-party script.
- Cookie-consent logic that prevents page content from loading for some visitors or bots.
- Infinite scroll that never exposes a paginated HTML route.
A useful worked example is a software pricing page. If the server response includes only “Plans” and the monthly prices arrive from an API after a region selector loads, an AI crawler may receive no prices, the wrong region’s price, or an incomplete plan comparison. The remedy may be server-side rendering, static generation, a progressively enhanced HTML table or a dedicated crawlable pricing route—not necessarily a wholesale framework change.
JavaScript-rendered content and LLMs: test the dependency
JavaScript is not inherently harmful to AEO or GEO. It powers valid interactions, personalization and rich applications. The question is whether the site’s central meaning still exists before JavaScript succeeds.
Google’s JavaScript SEO guidance provides a practical benchmark: use real URLs in links and make important content accessible to crawlers. Google can render JavaScript, but rendering involves more moving parts than parsing a response. A failed script, blocked API, slow third-party tag or hydration error can affect the rendered page even when the raw response returns HTTP 200.
For website visibility to LLMs, the safer operational rule is:
> Put the answer to the page’s main user question in the response HTML, then use JavaScript to enhance rather than create it.
That does not mean duplicating every interactive element in static markup. It means preserving the commercially important facts. On a service page, those facts might be the service name, geographic coverage, outcomes, constraints, proof points and contact path. On a product page, they may be the product name, price basis, key attributes, availability explanation and linked variants.
Teams should compare source and rendered DOM for at least 10 high-value URLs: the homepage, top service or category pages, pricing, a representative product, a location page, a key guide, contact, policy pages and any landing page that drives paid traffic. This sample is not statistically universal; it is a fast way to catch revenue-relevant gaps before expanding the crawl.
For a broader performance context, the comparison of Core Web Vitals audit tools, including Google, crawlers and Audra is relevant because long script chains and rendering delays can be a visibility risk as well as a user-experience problem.
Structured data, metadata and accessibility text
Schema.org provides a shared vocabulary for describing things such as organizations, products, articles, events, people and reviews. JSON-LD is often the least intrusive implementation because it can be placed in the page response without changing visible layout. However, structured data is supporting context, not a substitute for clear on-page content.
For example, a Product entity with name, description and offers can help machines interpret a product page. It cannot compensate when the visible page has no plain-language product description, the offer data is stale or the canonical product URL is blocked from crawling.
Metadata has a similar role:
- Title elements help identify the page topic.
- Meta descriptions are useful summaries but are not a reliable source of truth for every system.
- Canonical tags express a preferred URL when similar pages exist.
- Open Graph tags are mainly social-sharing metadata, not an AI visibility guarantee.
- Meta robots can restrict indexing behavior for crawlers that honor the directives.
Accessibility is a frequently overlooked part of this audit. Good alt text, labels, headings and transcript alternatives make content more understandable to people using assistive technology. They also make the page less dependent on visual interpretation alone. An image of a comparison chart with no text equivalent is difficult for many users and can leave crawlers with little reliable detail. Conversely, decorative image alt text stuffed with keywords is not a substitute for a genuine accessible description.
Audra’s combined audit approach is useful here because structured data, headings, accessibility problems and technical SEO issues should be reviewed in the same evidence set. For teams that need to interpret crawler findings rather than just count them, this guide to structured data markup items in a site audit provides related context.
robots.txt, llms.txt and Markdown routes are different tools
Three files and route patterns are often grouped together as “AI SEO.” They do different jobs.
robots.txt controls crawler requests, not AI answers
A site’s robots.txt file provides crawl instructions for user agents that choose to follow the Robots Exclusion Protocol. Google’s documentation is explicit that robots.txt is not a mechanism for keeping a page out of search results or securing confidential information. If sensitive content must not be public, it should require authentication or be removed from the public web.
For LLM visibility work, robots.txt is an access check. Teams should verify that important public paths are not unintentionally disallowed for relevant crawlers and that the file itself is reachable at the root domain. They should also recognize that a directive cannot guarantee that every actor behaves as intended.
llms.txt is a proposed discovery aid, not a visibility switch
The llms.txt proposal describes a Markdown file at a site root intended to offer a concise, machine-readable map of important pages and documentation. It can be useful for organizing canonical resources, especially for developer documentation, product help centers and large knowledge bases.
But llms.txt does not make inaccessible content readable. It does not force a crawler to render JavaScript, override robots.txt, earn an index position, add training data or produce an answer-engine citation. Adoption is not universal, so it should be treated as an optional supplemental route—not the first fix for an empty response HTML document.
Markdown routes can improve clarity, but only when maintained
A Markdown version of a guide, API reference or policy can give both people and machines a clean text representation. It is most useful when it is a maintained canonical or clearly related version of an existing resource, with stable URLs and links. A thin Markdown copy created solely for crawlers can create duplication, inconsistencies and maintenance debt.
A sensible order is: fix the page content and crawlability first; add structured data where it accurately represents the page; then consider llms.txt or Markdown routes where they improve resource discovery.
A practical LLM content visibility scanner workflow
An LLM content visibility scanner should not claim to predict every answer produced by ChatGPT, Claude, Gemini or Perplexity. It should expose observable conditions that affect whether important information can be discovered and understood.
A practical six-step audit looks like this:
- Choose priority URLs. Start with 10 to 25 pages tied to revenue, brand definition, high-intent information or customer support.
- Capture raw responses. Check status code, HTML body, title, headings, body copy, links, canonical tags, robots directives and JSON-LD before browser rendering.
- Compare the rendered DOM. Identify headings, copy, links, FAQs or tables present only after scripts execute.
- Inspect crawl directives. Review robots.txt, meta robots, canonicals, noindex patterns, redirects and XML sitemap inclusion.
- Assess legibility. Check semantic headings, descriptive links, alternative text, transcripts, page language and readable text equivalents for data visualizations.
- Test actual answer exposure separately. Use a defined set of prompts for relevant answer engines, record the date, locale, prompt wording, cited sources and result. Repeat over time rather than treating one response as permanent truth.
Audra is designed around the first five layers: local desktop checks for AI answer-engine visibility, technical SEO, performance, accessibility, best practices and links. This gives consultants and agencies a single local-first audit record instead of separate exports from multiple subscription tools. The sixth layer—what a specific answer engine says on a specific day—should be documented as an observation, not confused with crawler access.
For teams reconciling issue lists across tools, Semrush Site Audit issues versus Audra fix priorities can help frame why prioritization should focus on business-critical pages rather than raw issue totals.
Before-and-after examples: fixing the missing layer
Consider a fictitious B2B cybersecurity homepage. Before remediation, the initial HTML contains a logo, navigation labels and a JavaScript application shell. After hydration, the browser displays the H1 “Managed Detection and Response for Mid-Market Teams,” customer logos, product cards and a comparison table.
The audit finding is not merely “JavaScript detected.” The business finding is that the first response lacks the company’s core category, target customer and differentiation. A better implementation could server-render the H1, two-paragraph summary, primary navigation links and comparison-table text, while retaining JavaScript for demos, filters and animations.
A second example is a travel operator with a destination page. It has valid TouristTrip-style structured data but loads itinerary details, departure dates and inclusions only after a third-party booking widget responds. Here, JSON-LD may describe the offer, but it cannot replace the user-facing itinerary. Publishing a server-rendered itinerary summary and crawlable details page would improve both user resilience and machine readability.
In both cases, the before-and-after validation should include source HTML, rendered DOM, response status, internal links and accessibility text. It should also include a performance check: a page that technically renders but waits 12 seconds for a blocked API is not a robust visibility foundation.
What a crawler can fetch is not what an LLM will cite
This distinction prevents overclaiming. A public, fast, server-rendered page with accurate Schema.org markup may be easy to crawl and still never appear in a particular AI answer. Reasons can include query relevance, freshness, source reputation, retrieval design, geographic context, competing sources, licensing arrangements or the answer engine’s own product decisions.
Likewise, an answer engine may cite a page that has technical imperfections because it found the information through another route, a cached copy, a partner index or a browser-capable fetch. That outcome does not make the technical gap harmless.
AEO and GEO should therefore be measured in two connected but distinct dashboards:
- Eligibility and legibility: Can a crawler obtain a clear, permitted, semantically meaningful representation of the page?
- Observed exposure: For a controlled prompt set, is the brand or URL mentioned, accurately described or cited?
The first dashboard is where a website audit produces repeatable engineering work. The second is a changing market signal. Keeping them separate makes client reporting more honest and makes remediation more actionable.
Which should you choose: manual checks, crawlers or Audra?
Manual inspection is appropriate when a developer needs to diagnose one page immediately. Compare “View Page Source” with the browser’s inspected DOM, then test a representative API response. It is quick but hard to scale and easy to document inconsistently.
A browser-capable crawler is appropriate when the central question is Google-style JavaScript rendering across hundreds or thousands of URLs. It can reveal pages where visible content, links or headings appear only after rendering. The trade-off is that rendered crawling can be resource-intensive and still does not represent every LLM’s behavior.
A dedicated AI visibility scanner is appropriate when a team wants a focused review of raw HTML gaps, metadata, access and answer-engine exposure. Buyers should ask exactly which crawlers, models, locations and prompts the tool tests, because “LLM visibility” is not a single standardized metric.
Audra fits SEO consultants, agencies, marketers and site owners who need those checks alongside technical SEO, performance, accessibility and link auditing in a local desktop workflow. It is particularly suited to teams that need client-ready evidence without adding another recurring subscription. For Core Web Vitals interpretation, it also helps to understand why Google Search Console and Semrush Core Web Vitals results can differ, rather than assuming one measurement is universally definitive.
Verdict
Raw HTML is the baseline for robust website visibility to LLMs; rendered DOM shows what a browser can add; structured data and crawl directives add context and controls; llms.txt and Markdown routes can support discovery. None of these alone guarantees an AI citation.
The practical goal is simpler: make priority pages intelligible, accessible and crawlable in their initial response wherever possible, validate the rendered experience, then measure real answer-engine exposure as a separate changing signal. Audra brings those technical checks into one local-first audit so teams can show exactly what was found, what changed and what still needs testing.
FAQ
What content can LLMs see or miss on a website?
LLMs and their associated crawlers may be able to fetch raw HTML, links, headings, metadata and structured data, but their rendering capabilities vary. Content injected after JavaScript runs—such as tab panels, API-loaded copy or interactive product details—needs separate testing. A crawler fetching a page is also different from an answer engine choosing to cite it.
Do LLMs read JavaScript-rendered and dynamically loaded content?
Some systems can render JavaScript in some circumstances, while others may rely more heavily on the initial server response. There is no safe universal rule that every LLM ignores JavaScript or that every one renders it fully. Put business-critical text and links in response HTML where feasible, then compare source against rendered DOM on priority pages.
Can hidden HTML, developer comments or metadata affect what LLMs expose?
Potentially, yes. Content in the public response HTML can be fetched even if it is not prominent in the visual interface, so public developer comments and obsolete text deserve review. Metadata and structured data can provide machine-readable context, but they should accurately match visible page content. Confidential information should never be protected by obscurity or robots.txt alone.
Does llms.txt make a website more visible to AI search engines?
llms.txt can provide a concise Markdown map of important resources for tools that choose to use it. It does not guarantee crawling, rendering, indexing, training inclusion or citations, and it cannot repair a page whose key content is missing from the response HTML. Treat it as a supplemental discovery file after core content, links and crawl controls are correct.
How can a site owner audit whether a site is visible to LLMs?
Start with high-value URLs and compare raw response HTML with the browser-rendered DOM. Check headings, core copy, links, structured data, robots.txt, canonicals, accessibility text and performance blockers. Then run a consistent prompt set in relevant answer engines and log dates, locale, sources and outcomes. Audra can consolidate the technical audit layers into a client-ready local report.