# robots.txt is not an AI policy, but it is the only switch most crawlers honor

> llms.txt tells an assistant what's worth reading. robots.txt is the file that decides whether a crawler is allowed to fetch it. Here's how I split those jobs on this site.

Published Oct 1, 2026 by Joshua Oas · Tags: AI, SEO

A lot of the advice about AI crawlers treats `robots.txt` and `llms.txt` as the same kind of file. They are not. One is a permission list. The other is a map. Mixing them up is how sites end up blocking the assistants they wanted to help, or publishing a tidy index that nothing is allowed to fetch.

## What robots.txt actually controls

`robots.txt` lives at the root of your site and is a convention, not a security boundary. Well-behaved crawlers fetch it before they fetch anything else and skip paths you disallow. Badly behaved ones ignore it. It has never been a lock.

The useful part, for AI, is that the major labs publish user-agent names and say they honor the file:

- `GPTBot` is OpenAI's training crawler. `OAI-SearchBot` and `ChatGPT-User` are the ones that fetch pages while answering a question.
- `ClaudeBot` and `anthropic-ai` are Anthropic's crawlers. `Claude-User` is the fetch that happens during a chat.
- `Google-Extended` is Google's control for Gemini training, separate from `Googlebot`, which still indexes the site for Search.
- `PerplexityBot` and `CCBot` (Common Crawl) show up often enough to be worth an explicit line.

The split that matters is training versus inference. Blocking `GPTBot` does not stop ChatGPT from fetching a page when someone asks about you. Blocking `Googlebot` to "keep Gemini out" also pulls you out of Search. Those are different products wearing similar names.

## A file that says yes to answers and no to training

This is the policy I actually want: assistants may read the site while helping someone, training crawlers may not, and search engines may. That is expressible in the file, with the caveat that you are trusting each operator to use the user-agent they claim.

```text
User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: *
Allow: /

Sitemap: https://example.com/sitemap-index.xml
```

`User-agent: *` is the fallback. Specific groups win over it, so the training bots above are blocked and everything else, including `ChatGPT-User` and `Googlebot`, is allowed. I also point at the sitemap in the same file, because that is the one place every crawler already looks.

If you want the opposite policy, invert it. Disallow `ChatGPT-User` and `Claude-User` and your pages will not be fetched at answer time either. There is no third file that overrides this.

## What it does not do

- It does not hide a page. Anything linked from the public web can still be fetched by a browser, a script, or a crawler that ignores the file.
- It does not rank you. Search engines do not treat "I mentioned an AI bot" as a signal.
- It does not describe your content. That is `llms.txt`, and an assistant that is blocked here will never read it.
- User-agent strings change. The list above was right when I wrote it. Recheck the vendor pages before you copy it into production.

<!-- TODO: recheck OpenAI, Anthropic, Google and Perplexity user-agent docs before publishing -->

## Where the file lives

It has to be at the root: `https://example.com/robots.txt`. A file in a subdirectory is ignored. Crawlers request it by that exact path, before they request anything else.

On a static host, that is a text file in the public folder. On anything that generates routes, serve the same bytes from `/robots.txt` with `Content-Type: text/plain`. Do not put HTML around it, and do not redirect it. A redirect is one more request a crawler may not follow.

The policy is the only thing in the file. Paths I do not want indexed, like draft previews, are disallowed by name. Everything else is allowed, and the Markdown alternates stay reachable, because an assistant that is permitted to read the HTML should also be permitted to read the `.md`.

Keep it next to the rest of the site config, not buried in a CMS field you will forget. The failure I keep seeing is a production file that still lists a staging host, because someone edited a copy and deployed the other one.

## How to tell if it is working

Fetch it yourself (`curl -A GPTBot https://example.com/robots.txt` only shows you the file, not the decision). The decision shows up in your logs: a blocked bot requests `/robots.txt` and then stops, an allowed one keeps requesting pages. Server access logs are more honest than any dashboard the vendor gives you.

If a bot you blocked keeps requesting HTML, the file is doing its job and the bot is not. At that point the remedy is a firewall rule, not another line of text.
