OpenAI Crawlers robots.txt: Rules, Limits and Log Verification
A verification-first guide to setting OpenAI crawler rules, understanding each user agent’s scope, and checking whether the intended access policy works in server logs.
· 13 min read
A TechSEO discussion reviewed 72 company robots.txt files: 18 named GPTBot, 11 named OAI-SearchBot, and none named OAI-AdsBot. The practical risk is that OpenAI crawlers robots.txt rules aimed at one bot may leave separate OpenAI access paths unaddressed; this guide helps teams write an intentional policy and verify the result with live files and logs.
A Disallow rule for GPTBot is not a universal OpenAI block. OpenAI documents distinct user agents for automated model-training crawls, ChatGPT search crawling, user-triggered retrieval, and ad-related landing-page checks. For SEO consultants, agencies, and site owners, the payoff is a defensible crawler policy: the right rule for each goal, a standards-based configuration, and evidence of what actually reached the site.
Why GPTBot alone is not an OpenAI policy
The 72-site count came from a September 2026 TechSEO community discussion, not a controlled industry study. Its useful finding is narrower: of eight sites that blocked GPTBot, four did not name OAI-SearchBot. The discussion cannot establish whether those omissions were deliberate, accidental, or the most common policy among publishers.
That distinction matters because OpenAI’s crawler documentation describes separate agents with separate purposes. A site may choose to disallow GPTBot while allowing OAI-SearchBot, but that is one possible policy—not a default recommendation or a demonstrated industry norm.
A team should decide access by use case rather than applying an undifferentiated “block AI” rule:
- Decide whether the site permits automated GPTBot crawling.
- Decide whether pages should be crawlable by OAI-SearchBot for ChatGPT search features.
- Identify whether live user-requested retrieval by ChatGPT-User is acceptable for public pages.
- If ChatGPT ads are in use, make a separate decision for OAI-AdsBot.
- Document the owner, business rationale, deployment date, and validation evidence.
A broad User-agent: * directive can affect OpenAI agents, Googlebot, and any other bot that elects to follow it. That can be appropriate for a genuinely low-value public path such as /search/ or /staging/, but it is not a substitute for bot-specific decisions. Agencies should add crawler policy to the same remediation record as indexability, redirects, and performance issues; a prioritized SEO audit plan provides a useful model for recording impact and ownership.
The four OpenAI crawlers at a glance
OpenAI’s public bot documentation identifies four relevant user-agent tokens as of September 2, 2026. The table below separates documented purpose from the policy question a site owner needs to answer. It does not infer traffic volume, ranking impact, or the likelihood that any bot will visit a particular URL.
| User agent | Documented context | What the site owner decides |
|---|---|---|
GPTBot | Automated crawling associated with improving OpenAI models | Whether to permit this automated crawl path |
OAI-SearchBot | Crawling associated with search features in ChatGPT | Whether public pages may be crawled for this search-related use |
ChatGPT-User | Fetches made following a user action in ChatGPT | Whether public pages are suitable for live user-requested retrieval |
OAI-AdsBot | Crawling related to ChatGPT advertising and submitted landing pages | Whether relevant ad landing pages are accessible for this purpose |
GPTBot and OAI-SearchBot are separate automated policies
OpenAI describes GPTBot and OAI-SearchBot separately, so teams should configure and review them separately. Blocking GPTBot does not automatically create a block for OAI-SearchBot, and allowing OAI-SearchBot does not by itself create an allow rule for GPTBot.
For an SEO team, OAI-SearchBot is the agent most directly connected to ChatGPT search crawl eligibility. Allowing it is not a promise of a citation, referral, ranking, or favorable answer. It only removes a possible crawler-access barrier; retrieval and answer selection remain product decisions that vary by query and time.
ChatGPT-User and OAI-AdsBot need different treatment
ChatGPT-User is materially different from an automated discovery crawler because it is associated with an action initiated by a ChatGPT user. OpenAI’s documentation should be treated as the current authority on how that product handles robots instructions, rather than assuming GPTBot behavior applies unchanged.
OAI-AdsBot is relevant when a business submits landing pages for ChatGPT advertising. OpenAI’s advertiser guidance describes the bot and its access requirements. Teams not using that advertising product can still log and monitor the agent, but should avoid inventing a business need to allow it.
OpenAI crawlers robots.txt rules that follow RFC 9309
A robots.txt file is served at the host root—for example, https://www.example.com/robots.txt. RFC 9309 defines the Robots Exclusion Protocol syntax, including User-agent, Allow, and Disallow records.
The standard’s matching details matter. A crawler first selects the group with the most specific matching user-agent token. Within that selected group, the matching rule with the longest path value applies; if equally long Allow and Disallow rules match, Allow wins. A narrow allow therefore does not generally override a broader block merely because it is an allow—it must be at least as specific, and normally more specific, in the relevant group.
Pattern 1: block GPTBot and allow OAI-SearchBot
This pattern expresses one clear policy: do not permit GPTBot crawling, while permitting OAI-SearchBot access to public paths.
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
The configuration does not guarantee that ChatGPT will mention or cite the domain. It also does not state a policy for ChatGPT-User or OAI-AdsBot; those need explicit groups only if the site wants a distinct rule for them.
Pattern 2: allow both automated agents, restrict a directory
The following example permits both automated agents across the site but excludes /account/ from OAI-SearchBot. The longer Disallow: /account/ path is more specific than Allow: / in the same OAI-SearchBot group.
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
Disallow: /account/
A site that wants to re-allow one public subdirectory can use a still more specific path:
User-agent: OAI-SearchBot
Disallow: /account/
Allow: /account/public-help/
This is a robots convention, not access control. If /account/public-help/ should actually be private, it needs a login or server-side authorization rule—not a robots directive.
What robots.txt can and cannot guarantee
RFC 9309 provides a convention for crawler access; it does not make public content confidential. A URL protected only by robots.txt can still be requested by a browser, discovered through links, copied by another service, or visited by a client that does not follow the protocol.
For OpenAI, the safest approach is to separate documented controls from assumptions:
- GPTBot and OAI-SearchBot: OpenAI publishes bot-specific robots.txt guidance. Their rules should be reviewed as separate automated-crawl policies.
- ChatGPT-User: OpenAI documents this as user-triggered access. A site should not assume that a GPTBot rule creates an equivalent technical barrier for a user-directed fetch. Sensitive material requires authentication, authorization, or network restrictions.
- OAI-AdsBot: The relevant requirements should be checked against OpenAI’s advertiser crawler guidance, especially before diagnosing an inaccessible submitted landing page.
- Search visibility: An allowed crawler is an eligibility condition, not an inclusion guarantee. Content relevance, response status, rendering, internal linking, canonical signals, and the answer engine’s retrieval decisions can all affect outcomes.
Teams should also avoid treating robots.txt and noindex as interchangeable. robots.txt communicates crawling preferences. A noindex directive is an indexing instruction for crawlers that can fetch and process the page. Neither is a security control, and neither should protect previews, customer portals, private documents, or staging environments.
Googlebot needs its own review. A page can be available to Googlebot while unavailable to OAI-SearchBot, or the reverse, because the user-agent groups and product effects differ. This is particularly relevant when teams diagnose a situation where Google indexing is healthy but AI-search visibility is weak.
How to change a live robots.txt file safely
The frequent failure is not a typo; it is editing a local template while a CDN, CMS, reverse proxy, or separate hostname serves a different file. A complete review should inspect the public response that each relevant host actually returns.
Use this deployment sequence:
- Fetch
https://example.com/robots.txtand thewwwequivalent if both hosts respond. - Check relevant subdomains separately, such as
help.example.comorshop.example.com; each host has its own robots file. - Preserve existing
Sitemap:records and user-agent groups unless the change request explicitly replaces them. - Add exact tokens:
GPTBot,OAI-SearchBot,ChatGPT-User, andOAI-AdsBot. Do not use guessed names such asOpenAIBot. - Test the homepage, one editorial URL, one parameterized URL, one intended blocked URL, and one redirected URL.
- Save the prior file and deployment timestamp so later logs can be interpreted accurately.
Separate groups are easier to audit than a combined group when policies differ. For example, a single group listing two user agents applies the same records to both; it cannot express “block GPTBot but allow OAI-SearchBot.” Keeping the groups distinct reduces accidental policy coupling during future handoffs.
A path rule also applies only to paths that match it. Disallow: /internal-docs/ does not block a duplicate page published at /help/internal-docs/, a downloadable PDF in /files/, or content exposed on another host. Canonical URLs, redirects, and alternate publishing locations should be reviewed as part of the audit.
Verify real access in server, CDN and WAF logs
A robots.txt change is a configuration statement. Access logs provide evidence that a request occurred, while CDN and WAF logs reveal requests that may never reach the origin server.
A simplified access-log entry could look like this:
203.0.113.42 - - [02/Sep/2026:10:14:22 +0000] "GET /guides/technical-audit HTTP/1.1" 200 18472 "-" "... OAI-SearchBot ..."
The useful fields are the timestamp, request path, HTTP status, user-agent string, source IP, and any edge action. A 200 request for /robots.txt, followed by 200 responses for permitted pages, is stronger evidence than the robots file alone. A 403, 429, CAPTCHA, or managed challenge in Cloudflare, Akamai, Fastly, or a WAF log identifies a separate access problem.
A seven-day validation method
For a modest site, start with at least seven days of logs after a deployment; high-traffic sites may have enough evidence sooner, while low-traffic sites may need longer. No fixed period proves that a crawler will visit every eligible URL.
- Search separately for
GPTBot,OAI-SearchBot,ChatGPT-User, andOAI-AdsBot. - Record requests for
/robots.txt, requested page paths, statuses, and response sizes. - Segment denied events by CDN or WAF rule ID, not just origin status code.
- Compare observed paths with the intended policy and identify unexpected hostnames or redirects.
- Recheck after any cache purge, CDN configuration change, or robots deployment correction.
A user-agent header can be spoofed. OpenAI’s bot documentation publishes verification information, including links to current crawler IP data where applicable. Comparing source IPs with the current, bot-specific ranges published by OpenAI can increase confidence, but it is not definitive identity proof: infrastructure changes, proxy layers, incomplete log retention, and incorrect range handling can all produce false conclusions. The current OpenAI documentation, rather than a copied static IP list, should be the source used for each investigation.
A worked audit example for an agency
Consider an agency auditing www.example-client.com. The client wants ChatGPT search eligibility for its 120 public resource pages, does not want GPTBot automated crawling, and does not run ChatGPT ads.
The initial audit finds this live rule:
User-agent: GPTBot
Disallow: /
That rule records the GPTBot decision, but it says nothing specific about OAI-SearchBot. The agency first confirms the client’s intent, then deploys the two-group configuration from Pattern 1. It tests /robots.txt, /resources/seo-audit/, and /account/; the public resource returns 200, while the account area requires authentication.
The next log review finds OAI-SearchBot requesting /robots.txt with 200, but a request to a resource URL receives 403 at the CDN. The issue is not the robots rule. The agency investigates the edge security rule, creates an approved exception after the client’s security review, and records the before-and-after evidence.
This is where Audra can make the client deliverable more useful. An auditor can run Audra on the same representative public URLs to capture technical SEO, performance, accessibility, best-practice, link, and AI-visibility findings, then place the crawler-policy decision and log evidence beside those URL-level results in the client-ready report. A concrete finding might read: “Resource URL is crawl-permitted but returns a WAF 403 to the observed search bot; page also has two broken internal links.” That connects the access issue to the wider remediation work instead of presenting a robots.txt change as the whole AEO audit.
Measure AI-search outcomes separately from access
An OAI-SearchBot log entry does not prove a ChatGPT citation. Conversely, a lack of bot traffic in a short sample does not prove exclusion, because crawl scheduling and page selection are not published as a fixed frequency guarantee.
A useful monthly measurement model has three layers:
| Layer | Question | Evidence |
|---|---|---|
| Access | Can the intended agent retrieve intended public URLs? | robots response, log statuses, WAF events |
| Technical eligibility | Are pages usable after retrieval? | 200 status, canonical consistency, rendered content, links |
| Visibility outcome | Does the brand appear for relevant prompts? | documented prompt tests, cited URLs, referral analytics |
Keep a stable set of commercially relevant prompts and record the date, location, account state, and result format when testing. Pair that evidence with site health checks. For example, an inaccessible or orphaned page may fail to gain visibility for ordinary technical reasons, much as pages can be found but fail to progress through traditional search indexing; the diagnostic discipline in this guide to “Discovered – currently not indexed” is directly applicable.
For outcome assessment, teams should distinguish observed citations from broad claims that a brand is “winning AI search.” A practical AI-search evidence scorecard can help frame repeatable measurement. Crawler access establishes a prerequisite, not a performance result.
FAQ
Does robots.txt stop OpenAI’s AI crawlers?
OpenAI documents robots.txt controls for its named automated crawlers, including GPTBot and OAI-SearchBot. However, robots.txt is a voluntary crawl-access protocol, not a security barrier. ChatGPT-User is associated with user-triggered retrieval, so a site should not rely on a crawler directive to protect sensitive content. Use authentication or network controls for content that must not be fetched.
What are the four OpenAI crawlers and what does each one do?
As documented by OpenAI in September 2026, GPTBot is associated with automated model-improvement crawling, OAI-SearchBot with ChatGPT search features, ChatGPT-User with user-initiated retrieval, and OAI-AdsBot with ChatGPT advertising landing pages. Their presence as separate user agents means each should be reviewed as a separate access-policy decision.
What is the difference between GPTBot, OAI-SearchBot and ChatGPT-User?
GPTBot and OAI-SearchBot are distinct automated agents with different documented product contexts. ChatGPT-User is tied to a user action, rather than normal automated discovery crawling. Therefore, blocking GPTBot does not automatically block OAI-SearchBot, and neither rule should be assumed to create an equivalent protection against a user-requested fetch.
How should robots.txt rules be written for OpenAI crawlers?
Use each exact user-agent token in its own group when policies differ. Under RFC 9309, use Allow and Disallow paths carefully: the longest matching path rule applies, with Allow winning only on an equal-length tie. Fetch the live robots.txt file after deployment and test representative URLs, hosts, redirects, and restricted paths.
How can a site tell which OpenAI bot is visiting it?
Search origin, CDN, and WAF logs for GPTBot, OAI-SearchBot, ChatGPT-User, and OAI-AdsBot, then inspect timestamps, paths, statuses, and edge actions. User-agent strings can be spoofed. Compare source IPs with the current bot-specific verification information published in OpenAI’s documentation, but treat that comparison as supporting evidence rather than definitive proof.