Which AI crawlers to allow, and how

One small file decides whether AI tools can read your website at all. Here is what each crawler does, and how to check yours in two minutes.

Last updated: 19 September 2026

The short version

Your site has a file at yoursite.com/robots.txt. It tells automated visitors which pages they may fetch. AI companies send their own visitors, and each one has a name.

If that file blocks them, assistants cannot read your pages. They can still name you from other sources, like a directory or a review site, but your own words are out of the conversation. That is a bad trade for a business trying to get hired.

The four names that matter

These are the ones we check on every scan.

  • GPTBotis OpenAI’s crawler. It collects pages that help ChatGPT learn what exists.
  • OAI-SearchBot is the one that fetches pages when ChatGPT is answering a question with search. This is the one that affects whether you get named today.
  • PerplexityBot does the same job for Perplexity.
  • Google-Extended does not control normal Google search. It controls whether Google may use your pages for Gemini and its AI answers.

Blocking Google-Extended is the one people get wrong most often. They think they are protecting their search ranking. They are not. They are opting out of AI answers while leaving normal search exactly as it was.

Check yours in two minutes

Open yoursite.com/robots.txt in a browser. You are looking for any block that names one of the four crawlers above, or one that names *, which means everyone.

This is what a block looks like. If you see this, an assistant is being turned away.

robots.txt - blocking, probably not what you want
User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Disallow: / means the whole site. A line reading Disallow: with nothing after it means the opposite: nothing is blocked.

What to put there instead

For most local businesses the right answer is to let them all in, and keep out only the pages that were never for the public. Checkout pages, admin screens, and internal search results.

robots.txt - a sensible starting point
User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /admin

Sitemap: https://YOURSITE.com/sitemap.xml

One rule for everyone, a short list of private paths, and a line pointing at your sitemap. You do not need to name the AI crawlers to allow them. Not blocking them is the same as allowing them.

Two things this file cannot do

It cannot make an assistant recommend you. Letting a crawler in is permission, not persuasion. What gets you named is having pages worth quoting, reviews worth citing, and listings that agree with each other.

It also cannot hide a page from the public. Anyone who has the link can still open it. If a page must stay private, it needs a password, not a line in this file.

Common questions

What happens if I block GPTBot?

ChatGPT stops using your pages to train and, depending on the bot, stops reading them when it answers a question. Blocking it does not remove you from the internet. It removes one way for an assistant to learn what you sell.

Is blocking AI crawlers ever the right call?

For a publisher whose words are the product, sometimes. For a local business trying to get hired, almost never. You want to be quoted.

Does allowing these crawlers slow my site down?

No. They fetch pages the same way a search engine does, a few at a time. If your site can handle Google, it can handle these.

I did not write my robots.txt file. Who did?

Usually your website platform or a plugin. Wix, Squarespace, Shopify and most WordPress SEO plugins generate one for you, and some ship with AI crawlers blocked by default.