robots.txt for the AI Era: Which Crawlers to Actually Let In
By Andrew Pyle
Allow OAI-SearchBot and ChatGPT-User if you want to show up in ChatGPT search and be readable when someone asks about your page. Block GPTBot only if you specifically object to your content training a model, and treat that as a separate decision from the first one. That is the whole answer. The rest of this is why, and how to write the actual rules.
I already wrote about AEO and GEO as strategy: be specific, be structured, be genuinely expert. This is not that essay. This is the tactical one: which named bots exist, what each one is for, and the exact robots.txt lines to write so you get found without a rule you did not mean to write.
01
Four bots, four different jobs
People talk about "AI crawlers" like there is one. OpenAI alone runs four, and they do not do the same thing.
- **OAI-SearchBot** crawls pages to build the search index that ChatGPT's search feature draws on. This is the one that gets you found when someone asks ChatGPT a question and it searches the live web for an answer.
- **ChatGPT-User** fetches a specific page in real time, mid-conversation, when a user asks ChatGPT about that exact URL or when a plugin/browsing action needs the page right now. It is a live fetch, not an index-builder.
- **GPTBot** collects pages for training data. This has nothing to do with whether you show up in a ChatGPT answer today. It is about whether your text becomes part of a future model's weights.
- **OAI-AdsBot** validates ads, a narrower job most sites never need to think about.
Four names, four purposes. Treating them as one line item, "AI bots," is where the mistakes start.
02
The mistake: one blanket rule that blocks everything
The costly error I see is a robots.txt that disallows every user-agent matching some AI pattern, in one sweeping block. It usually reads like a copy-pasted "protect yourself from AI" template, and it catches OAI-SearchBot in the same net as GPTBot.
That single rule does two things, and only one of them is what the site owner wanted. It keeps GPTBot from training on the content, which was probably the intent. It also removes the site from ChatGPT search entirely, which was probably not the intent, and which protects nothing that was already trained. If GPTBot crawled you last month, blocking it today does not un-train anything. You just lose the upside (being found) while keeping none of the protection you thought you were adding.
I have found this exact rule sitting in inherited robots.txt files that nobody had actually reviewed since a "just block the AI stuff" pass. The line looks defensive. It is actually a self-inflicted delisting from a channel that is only going to matter more.
03
Two decisions, not one
The fix is to stop treating "AI crawlers" as a single yes/no and make two separate calls.
**Decision one: do I want to be findable in AI search?** If yes, allow OAI-SearchBot and ChatGPT-User. There is no version of "be visible in ChatGPT answers" that works with those two blocked.
**Decision two: do I want my content used to train models?** This is independent of decision one. You can allow OAI-SearchBot for visibility and still block GPTBot for training, and that combination is coherent, not contradictory. It says: cite me, but do not train on me. Plenty of publishers land exactly there.
A minimal robots.txt that makes both calls explicitly:
``` User-agent: OAI-SearchBot Allow: /
User-agent: ChatGPT-User Allow: /
User-agent: GPTBot Disallow: /
User-agent: OAI-AdsBot Disallow: / ```
That is four short blocks instead of one blanket rule, and every line is a decision you actually made, not one you inherited.
04
Verify by IP, not by user-agent string
robots.txt is a request, not a lock. Anything can claim to be GPTBot in its user-agent header, including a scraper you would rather not have. If you are doing anything more than the honor-system robots.txt rule, such as server-side blocking by user-agent, verify the request is genuinely OpenAI's before you trust it.
OpenAI publishes the IP ranges each bot uses, at `openai.com/searchbot.json`, `openai.com/chatgpt-user.json`, and `openai.com/gptbot.json`. A request claiming to be OAI-SearchBot from an IP not in that published range is not OAI-SearchBot. That distinction matters if you are logging crawler behavior, gating content by bot identity at the edge, or trying to figure out why a page got scraped in a way that does not match what OpenAI says it does. Check the header against the published range before you act on either one.
05
Bing feeds ChatGPT search more than Google does
This is the part that trips up people who only ever set up Google Search Console. Bing's index is the documented, primary source behind ChatGPT search. If your page is indexed in Google but was never submitted to Bing, you are invisible to the index ChatGPT has leaned on most. In 2026, ChatGPT also crawls directly through OAI-SearchBot, and independent testing suggests it sometimes pulls from Google's index too, but Bing is the one piece of that stack you can control directly and for free.
The action is small and gets skipped constantly: sign up for Bing Webmaster Tools, verify the site, submit the sitemap. It takes fifteen minutes and most sites that only optimize for Google have simply never done it.
06
Keep the file honest
One more failure mode, and it has nothing to do with AI specifically: robots.txt rules stack, and a broad `Disallow` written for an unrelated reason can quietly catch a bot you meant to allow. If you have a rule blocking `/api/` or `/search/` for load reasons, and OAI-SearchBot happens to need a path under one of those, the more specific `Allow` block for that user-agent has to come after or the crawler may still be shut out depending on how your server evaluates the file. Read the whole file top to bottom as one document before you ship it, not rule by rule.
07
The decision checklist
Before you touch robots.txt for AI crawlers, answer these in order:
- [ ] Do I want to be findable in ChatGPT search and other AI answers? If yes, allow OAI-SearchBot and ChatGPT-User explicitly, by name.
- [ ] Do I want my content used as AI training data? Answer this separately from the question above. If no, disallow GPTBot specifically. Do not use a wildcard that also catches OAI-SearchBot.
- [ ] Am I indexed in Bing, not just Google? If not, submit the sitemap to Bing Webmaster Tools this week.
- [ ] Do I need to verify crawler identity anywhere beyond robots.txt, such as server logs or edge rules? If so, check requests against OpenAI's published IP ranges, not the user-agent string alone.
- [ ] Have I read the full robots.txt file as one document, checking that no older, broader rule accidentally blocks a bot I just allowed?
08
The bottom line
The AI-crawler question is not "block or allow." It is two separate questions wearing one trench coat: do you want to be found, and do you want to be trained on. Answer them separately, name the bots explicitly instead of pattern-matching "AI" as one category, and verify by IP if you are doing anything beyond the honor system. The single most expensive mistake here is a well-intentioned blanket block that takes out OAI-SearchBot along with GPTBot, because that trades a channel that is only growing for a protection that was never real to begin with, since a bot that already crawled you already has what it came for.
---
Written by Andrew Pyle. I run robots.txt and crawler policy across my own portfolio of sites and check it against OpenAI's published bot documentation and IP ranges directly, not secondhand summaries. [Operator: add one specific, true credential or result here before publishing, per E-E-A-T. Do not invent one.]
Related
writing
Answer Engine Optimization: A Builder's Guide to AEO
How to win the answer, not just the ranking
writing
A billion indexed pages is a direction, not a page count
I keep an absurdly large number as a north star — but the number isn't a target to hit by publishing a billion things. It's a forcing function that makes you solve the problems that only show up at scale, and it fails the moment you treat page count as the point.
writing
A sitemap should only list pages you actually want found
A sitemap is a set of recommendations you make to a search engine. Listing pages you've told it not to index, or that redirect elsewhere, is contradicting yourself — and a sitemap that contradicts itself teaches the crawler to trust you less.
writing
AEO and GEO: optimizing for answer engines
Two acronyms showed up in every SEO conversation this year. Here is the plain, first-hand version of what they mean and what actually earns them.
writing
I let agents run my SEO
The same command center that builds my sites now runs their SEO — as reversible, human-gated agent tasks. And the target it aims at has quietly moved.