· 3 min read · Joshua Oas
robots.txt is not an AI policy, but it is the only switch most crawlers honor
- AI
- SEO
A lot of the advice about AI crawlers treats robots.txt and llms.txt as the same kind of file. They are not. One is a permission list. The other is a map. Mixing them up is how sites end up blocking the assistants they wanted to help, or publishing a tidy index that nothing is allowed to fetch.
What robots.txt actually controls
robots.txt lives at the root of your site and is a convention, not a security boundary. Well-behaved crawlers fetch it before they fetch anything else and skip paths you disallow. Badly behaved ones ignore it. It has never been a lock.
The useful part, for AI, is that the major labs publish user-agent names and say they honor the file:
GPTBotis OpenAI’s training crawler.OAI-SearchBotandChatGPT-Userare the ones that fetch pages while answering a question.ClaudeBotandanthropic-aiare Anthropic’s crawlers.Claude-Useris the fetch that happens during a chat.Google-Extendedis Google’s control for Gemini training, separate fromGooglebot, which still indexes the site for Search.PerplexityBotandCCBot(Common Crawl) show up often enough to be worth an explicit line.
The split that matters is training versus inference. Blocking GPTBot does not stop ChatGPT from fetching a page when someone asks about you. Blocking Googlebot to “keep Gemini out” also pulls you out of Search. Those are different products wearing similar names.
A file that says yes to answers and no to training
This is the policy I actually want: assistants may read the site while helping someone, training crawlers may not, and search engines may. That is expressible in the file, with the caveat that you are trusting each operator to use the user-agent they claim.
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap-index.xml
User-agent: * is the fallback. Specific groups win over it, so the training bots above are blocked and everything else, including ChatGPT-User and Googlebot, is allowed. I also point at the sitemap in the same file, because that is the one place every crawler already looks.
If you want the opposite policy, invert it. Disallow ChatGPT-User and Claude-User and your pages will not be fetched at answer time either. There is no third file that overrides this.
What it does not do
- It does not hide a page. Anything linked from the public web can still be fetched by a browser, a script, or a crawler that ignores the file.
- It does not rank you. Search engines do not treat “I mentioned an AI bot” as a signal.
- It does not describe your content. That is
llms.txt, and an assistant that is blocked here will never read it. - User-agent strings change. The list above was right when I wrote it. Recheck the vendor pages before you copy it into production.
Where the file lives
It has to be at the root: https://example.com/robots.txt. A file in a subdirectory is ignored. Crawlers request it by that exact path, before they request anything else.
On a static host, that is a text file in the public folder. On anything that generates routes, serve the same bytes from /robots.txt with Content-Type: text/plain. Do not put HTML around it, and do not redirect it. A redirect is one more request a crawler may not follow.
The policy is the only thing in the file. Paths I do not want indexed, like draft previews, are disallowed by name. Everything else is allowed, and the Markdown alternates stay reachable, because an assistant that is permitted to read the HTML should also be permitted to read the .md.
Keep it next to the rest of the site config, not buried in a CMS field you will forget. The failure I keep seeing is a production file that still lists a staging host, because someone edited a copy and deployed the other one.
How to tell if it is working
Fetch it yourself (curl -A GPTBot https://example.com/robots.txt only shows you the file, not the decision). The decision shows up in your logs: a blocked bot requests /robots.txt and then stops, an allowed one keeps requesting pages. Server access logs are more honest than any dashboard the vendor gives you.
If a bot you blocked keeps requesting HTML, the file is doing its job and the bot is not. At that point the remedy is a firewall rule, not another line of text.