Skip to main content


GPTBot, ClaudeBot and PerplexityBot: How to Control AI Crawlers

GPTBot, ClaudeBot and PerplexityBot: Controlling AI crawlers with robots.txt by allowing or blocking each user agent
Dominik Breitbach

Dominik Breitbach · Founder & Lead SEO Strategist at taismo

Dominik is the founder and managing director of taismo, an SEO and GEO agency from Munich. He has worked in search marketing since 2010 and focuses on ongoing SEO support and visibility in AI answers, for companies in Germany and for international firms that want to be found in the German market.

⏱ Reading time: 14 min
🔄 Last updated: 1 October 2026

GPTBot is the web crawler OpenAI uses to collect content for training its AI models, and you control it with a User-agent: GPTBot group in your robots.txt. Blocking GPTBot keeps your pages out of future OpenAI training data; it leaves your citations in ChatGPT search untouched, because those come from a second crawler called OAI-SearchBot. That split runs through the whole industry: AI crawlers do three different jobs, and each job has its own user agent. This guide lists the 17 user agents that matter in 2026, shows how robots.txt rules are evaluated, where they stop working, and how to decide what to block.

👉 Before you change a single line, see how AI systems read your page today: Free AI visibility checker

AI crawlers do three jobs: Training, search index and live fetch

The question “should we allow AI crawlers?” has no single answer, because it bundles three processes that work very differently. The general concept is covered in our glossary entry on the AI crawler. This guide is about control: Which user agent to address, with which rule, and what the rule changes.

AI crawlers fall into three roles:

  1. Training. The bot collects text that a model later learns from. Whatever it takes ends up in the model weights and cannot be pulled back out. Training alone gives you no citation and no link.
  2. Search index. The bot builds a separate index that the AI system cites from when a user asks a question. This is where your brand gets named together with a link to your page. The search index role is the one that drives your visibility in AI answers.
  3. Live fetch. A person just asked a question, and the system loads your page at that moment. Technically this is a user-triggered request and no crawl, which is why several of these agents openly ignore robots.txt.

The detail that costs companies visibility: One operator runs several bots with different roles, and each bot has its own user agent. OpenAI runs four, Anthropic three, Perplexity two. Anyone who blocks only the best-known name, usually GPTBot, almost always targets the wrong job.

The three jobs of AI crawlers and their consequencesThree jobs, three consequencesOne operator often runs all three, each with its own user agent.1 · TrainingText flows into the model.No source credit,cannot be taken back.GPTBot · ClaudeBotGoogle-Extended · CCBot2 · Search indexThis is where the mentionand the link to your pagecome from. Visibility.OAI-SearchBotPerplexityBot · Claude-SearchBot3 · Live fetchA person just asked.robots.txt usuallydoes not apply here.ChatGPT-UserPerplexity-User · Claude-UserBlocking all three with one line also hits job 2.Job 2 is where AI answers name you. Job 3 ignores that line anyway.Fig. 1 · taismo
Fig. 1: The three jobs of AI crawlers. A blanket block on all three also removes the job that creates mentions in AI answers.

The AI crawler list 2026: User agents, operators and purpose

The table below was checked against each operator’s own documentation, because second-hand lists go stale within months. The first column is exactly what goes into the User-agent line. Case does not matter there: Google’s robots.txt specification treats the user agent token as case-insensitive.

User agent Operator Role Honors robots.txt
GPTBot OpenAI Training yes
OAI-SearchBot OpenAI Search index (ChatGPT search) yes
ChatGPT-User OpenAI Live fetch may not apply
OAI-AdsBot OpenAI Checks landing pages of ChatGPT ads yes
ClaudeBot Anthropic Training yes
Claude-SearchBot Anthropic Search index yes
Claude-User Anthropic Live fetch yes
Google-Extended Google Gemini training and grounding yes
Googlebot Google Google Search and AI Overviews yes
PerplexityBot Perplexity Search index yes
Perplexity-User Perplexity Live fetch generally ignored
meta-externalagent Meta Training yes
meta-webindexer Meta Search index yes
meta-externalfetcher Meta Live fetch and agents may not apply
Applebot-Extended Apple Usage signal for training yes
Applebot Apple Siri, Spotlight, Safari yes
CCBot Common Crawl Open web archive, widely used for training yes

Three rows deserve a second look. OAI-AdsBot is new: According to OpenAI, it validates the safety of web pages submitted as ads on ChatGPT and only visits those submitted landing pages. Applebot-Extended does no crawling at all. The token only signals how content that Applebot already collected may be used for training. And CCBot belongs to the nonprofit Common Crawl Foundation, no AI company. Its open archive is still one of the most widely used training sources, so a block there indirectly reaches many models.

OpenAI, Anthropic, Perplexity and Common Crawl publish the IP ranges of their bots as JSON files. Matching a request’s IP address against these files is the only reliable way to tell a real bot from a fake one; the file locations are listed in the sources below.

GPTBot vs. OAI-SearchBot: Two switches that get mixed up

The most expensive misconception in this field sounds like this: “We blocked GPTBot, so we are out of ChatGPT.” That statement is wrong in both directions.

GPTBot controls training, OAI-SearchBot controls the search index. OpenAI states that “each setting is independent of the others”. If you block GPTBot and allow OAI-SearchBot, ChatGPT search keeps citing you, and your texts stay out of model training. If you block only OAI-SearchBot, you disappear from ChatGPT’s search answers and still feed the training set as long as GPTBot gets through. The ChatGPT side of this is covered in depth in our guide to ChatGPT SEO.

Google draws the same line between Googlebot and Google-Extended. Google writes that Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” One detail is often missed: According to the same documentation, Google-Extended governs training of future Gemini models and grounding, which means feeding Search content to Gemini Apps and Vertex AI at answer time. A Google-Extended block therefore reaches further than training. Apple describes Applebot-Extended in the same spirit: Pages that block it can still appear in Apple’s search features.

In practice this means: A training block is a copyright decision. A search index block is a visibility decision. Writing both into the same line merges two decisions without looking at either. How AI systems choose which sources to name is explained in our article on ranking in AI answers.

💡
Pro tip: Before you change a line, write down the role of every user agent you plan to block. If the column says “search index”, you are about to remove your own mentions from AI answers. That one column prevents the most common wrong decision in this field.

How crawlers read your robots.txt: Four rules

A robots.txt file looks simpler than it is. The format is standardized as the Robots Exclusion Protocol in RFC 9309 (2022), and four rules decide whether your file does what you think it does.

Rule 1: A crawler follows exactly one group. It picks the group with the most specific matching user agent and ignores all others. If User-agent: * contains a rule and a separate group for GPTBot follows further down, GPTBot obeys only its own group. The wildcard rules are not added on top. This is how a newly added bot group can silently cancel an existing block.

Rule 2: When rules conflict, the longer path wins. Google measures the length of the rule path; with equal length, the least restrictive rule applies. That lets you open a single directory while the rest of the site stays closed.

Rule 3: A server error on robots.txt halts crawling. A 404 means “no restrictions” to Google. A 5xx error makes Google stop crawling the site for the first 12 hours; after that, Google uses the last good copy for up to 30 days. A robots.txt that throws a 503 under load is more dangerous than no file at all.

Rule 4: Disallow does not mean noindex. Google says a blocked URL “might still be indexed without visiting the page” if other pages link to it, and then appears without a description. To keep a page out of the index, you need a meta robots tag or the X-Robots-Tag HTTP header, and the page has to stay crawlable so the crawler can read that instruction. Setting both at once achieves the opposite: The bot may not load the page, so it never sees your noindex.

How a crawler evaluates robots.txtFour steps before a rule applies1. Pick a groupThe most specificuser agent wins.Only ONE group applies.2. Longest ruleMore characters in the pathbeat a shorter path.Same length: Allow.3. File reachable?404 = no rules.5xx = crawling stops,then up to 30 days cache.4. RuleappliesAnd the most common mix-up: Disallow does not mean noindexDisallowThe content is not read.The URL can still end up in the index.noindexThe URL stays out of the index.Only read if the page stays crawlable.Fig. 2 · taismo
Fig. 2: How a crawler evaluates robots.txt. Only one group applies, the longest rule wins, and Disallow alone does not prevent indexing.

Here is an example that sets the two switches separately: Training crawlers blocked, search crawlers and live fetchers open.

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

Sitemap: https://www.example.com/sitemap.xml

This file does not touch Googlebot, and that is deliberate. Googlebot feeds classic Google Search and AI Overviews. Blocking it costs you far more than a slice of AI visibility: It takes you out of Google Search entirely.

Where robots.txt reaches its limits

Three gaps remain, and they explain why a clean robots.txt is only half of the answer.

Live fetch agents follow their own rules, and their operators say so. OpenAI writes about ChatGPT-User that “robots.txt rules may not apply”, because a user initiated the request. Perplexity states that Perplexity-User “generally ignores robots.txt rules”. Meta says the same about meta-externalfetcher. When a person asks an AI system about your page, the page gets loaded, whatever your file says. Anthropic is the exception: According to its documentation, Claude-User lets site owners control which sites these user-initiated requests can access.

A user agent can be faked. The user agent is free text in the request header. A server rule that only checks this text blocks the bots that identify themselves honestly and lets every serious scraper through. Server-side blocking needs an IP check against the published ranges.

Third parties block for you without telling you. On July 1, 2025, Cloudflare announced it was changing its default to block AI crawlers. Add bot protection that answers fast request bursts with a 429, and hosting providers that filter at network level. The result is a contradiction nobody sees in the source code: The robots.txt file says Allow, the server door says 403.

Cloudflare backed the change with a striking comparison: Getting traffic from OpenAI is 750 times harder than it used to be from Google, and from Anthropic 30,000 times harder. AI systems read a lot and send back little. Concluding that allowing them does not pay mixes up two metrics. The value of a mention in an AI answer lies in the shortlist it puts you on, long before any click. That is why answer engine optimization (AEO, also called generative engine optimization or GEO) is measured in mentions and citations first and in sessions second.

The legal side matters for anyone selling into Europe. Article 4 of the EU Directive on Copyright in the Digital Single Market lets rights holders reserve their content from text and data mining, and for content published online it names “machine-readable means” as the appropriate form. Germany implemented this in Section 44b of its Copyright Act, and the EU AI Act obliges providers of general-purpose AI models to identify and comply with such reservations. robots.txt is the most common machine-readable way to declare that reservation. Whether a specific block holds up in a dispute is a legal question for a lawyer. In the United States there is no comparable statutory opt-out; robots.txt there works as an industry convention that the large operators commit to in their documentation.

Do AI systems actually reach your website?

Our AEO audit checks exactly that: Which user agents arrive, which status codes they get, whether robots.txt and server agree, and where visibility in AI answers is lost. You get findings in priority order, ready for your team or for us to implement.

See the audit

Block or allow GPTBot: One question decides

We sell visibility in AI answers, so here is our disclosure up front: We have an interest in companies allowing search crawlers. There are still cases in which a block is the right call, and one question separates them.

Is your content the product or the advertising for the product?

If people pay for your texts because the texts themselves are the value, every reuse in an AI answer costs you revenue. That applies to publishers, course providers, databases, image archives and paywalled newsrooms. For them, a training block is the obvious choice, and often a search index block as well.

If your texts explain what you sell, the math flips. A service page, a guide or a glossary has one job: To be found and understood. Blocking here removes you from exactly the questions in which a buyer looks for a provider, and protects texts you give away anyway. That describes most companies with a service that needs explaining, from B2B manufacturers to software vendors.

Block or allow AI crawlers: The decision in one questionOne question decidesIs your content the product, or the advertising for the product?The content IS the productPublisher, course provider, database,image archive, paywalled newsroomBlock trainingWeigh the search index, often block it too.Every reuse costs revenue here.The content PROMOTES the productService page, guide, glossary,a service that needs explainingAllow the search indexDecide on training separately.Blocking costs you the mention.In both cases: Deciding per directory is cleaner than all or nothing.Fig. 3 · taismo
Fig. 3: Block or allow AI crawlers. The decision follows from one question about the role of your content.

The middle ground we recommend in most cases

Block training crawlers, allow search crawlers. You keep the mention with a link and declare your text and data mining reservation at the same time. This option has a price, and it belongs on the table: If a model does not know your brand from training, it names you less often when it answers without a live web search. For brands with little awareness, allowing training on purpose can make sense. How the two disciplines interact is covered in our comparison of AEO and SEO.

A third option is rarely mentioned and often fits best: Decide per directory. The guide section stays open, the member area, download folder or paid archive is blocked. Because the longer rule wins, this fits cleanly into a single group:

User-agent: GPTBot
Disallow: /members/
Disallow: /downloads/
Allow: /

Server load and other weak arguments

Server load is rarely the deciding factor. If a bot really fetches too much, the right answer is a Crawl-delay directive or rate limiting. Anthropic explicitly supports the non-standard Crawl-delay extension and shows it in its own documentation. Crawl volume becomes a crawl budget topic for very large sites; the typical company website with a few hundred URLs stays far below that line.

“Everyone else blocks it” is just as weak. What is right for a newspaper is wrong for an engineering firm, and the other way round. The user agents are the same for everyone; the decision is individual.

💡
Pro tip: Put a recurring reminder at the start of each quarter: Compare your robots.txt with the current operator documentation, then break down your access logs by user agent. Fifteen minutes catch new bots before they run unnoticed for a year.

How we configure our own robots.txt

Anyone giving advice should show their own setup. Our robots.txt is short on purpose. It blocks one thing, the paginated archive pages under /page/*, and points to the sitemap. We block no AI crawler at all, neither for training nor for search. That is a decision: Our texts are the advertising for our service.

The positive counterpart is our llms.txt, a briefing document for language models that we maintain by hand. It currently has 19,136 characters with 94 annotated links, bilingual in one file. It tells a model in structured form who we are, which claims are backed by evidence and which page answers which question. What the format can and cannot do is explained in our glossary entry on the llms.txt file; the way from draft to launch is described in our llms.txt guide.

Two lessons from running it that you will not find in the documentation:

  1. A static root file deserves a deployment with backup, health check and rollback. Our script saves the live version, uploads the new one, checks that it loads and rolls back automatically on failure. When you compare your local copy with the live file, preserve the line endings. Otherwise the comparison reports a difference that does not exist, which cost us one wrong diagnosis.
  2. A language folder can be virtual. We wanted a separate llms.txt for our English site under /en/. In our multilingual setup, /en/ is a rewrite route, and the rewrite rule only fires while no real folder of that name exists. Uploading en/llms.txt would have created the folder and taken the English homepage offline. The clean solution is a virtual route in code.

And one classification that has to be honest: The llms.txt file is an offer to AI systems. Jeremy Howard proposed the format in September 2024. No operator has committed to reading it, and it controls nothing. Access is governed by robots.txt and the server, and llms.txt complements both.

How to verify that your AI crawler setup works

The most common state in practice is an unnoticed decision. robots.txt says one thing, the server does another. Four steps settle this in fifteen minutes:

  1. Step 1: Fetch your robots.txt yourself and check the status code. You want a 200. A 5xx is an emergency, a 404 means no rules apply at all.
  2. Step 2: Request your page as a bot. This separates what you configured from what actually happens. A request with the user agent set shows the real status code.
    curl -sI -A "GPTBot" https://www.example.com/
    curl -sI -A "OAI-SearchBot" https://www.example.com/
    curl -sI -A "ClaudeBot" https://www.example.com/
    curl -sI -A "PerplexityBot" https://www.example.com/

    A 403, a 429 or a redirect to a challenge page means a firewall, bot protection or your host is blocking. That block appears in no robots.txt and in no normal crawl, because nobody asks as a bot.

  3. Step 3: Read your access logs. Set up a report by user agent and watch for four weeks which agents arrive, how often, and with which status code. A block you set has to show up as a 403. A bot that mostly sees 404s has a site structure problem and no access problem.
  4. Step 4: Check suspicious traffic against the IP lists. If a user agent shows up unusually often, match its IP with the operator’s JSON file. An address missing from the file belongs to someone using the name. That tells you whether to block a user agent or a single address.

Five robots.txt mistakes we see in audits

These five come up again and again, across industries and content management systems:

  • Everything blocked because a template said so. Ready-made AI blocklists are mostly written by publishers for publishers. A service company that copies one loses exactly the visibility it pays for elsewhere.
  • Disallow and noindex on the same page. The page stays in the index without a snippet, the opposite of the intent.
  • A new bot group cancels the wildcard rules. Since only one group applies, a bot with its own group loses all general blocks. Block /internal/ under User-agent: *, add a GPTBot group later, and /internal/ is open to GPTBot.
  • Server-side blocking without an IP check. A rule that only reads the user agent text hits the honest bots and misses everyone else.
  • The file dates from 2023 and nobody looks at it. OAI-SearchBot, Claude-SearchBot, meta-webindexer and OAI-AdsBot did not exist when the debate started. robots.txt belongs on the table once a quarter, together with the logs.

None of these mistakes is expensive to fix. They are expensive because they stay unnoticed: Nobody reads the file after it was created. If you would rather not keep an eye on it yourself, the quarterly check fits into ongoing SEO support, where it belongs anyway.

Let’s look at your setup together

Not sure whether your robots.txt does what it should? We clarify that in a short call, in English or German. We tell you what we see, and you decide whether to implement it yourself or hand it over.

Book a free first call

FAQ about GPTBot and AI crawlers

What is GPTBot?

GPTBot is OpenAI’s web crawler that collects publicly available content for training its AI models. You can block it with a User-agent: GPTBot group in robots.txt, and OpenAI publishes its IP ranges for verification.

Should I block GPTBot?

Block GPTBot if your content is your product, for example paid articles, courses or databases. If your content promotes your services, allowing it is usually the better choice, and in both cases you should keep OAI-SearchBot open for ChatGPT search.

Does blocking GPTBot remove my website from ChatGPT?

Blocking GPTBot keeps your content out of OpenAI’s model training; your citations in ChatGPT search stay, because they depend on OAI-SearchBot. Only blocking OAI-SearchBot removes you from ChatGPT’s search answers.

Do AI crawlers respect robots.txt?

The training and search crawlers of the large operators state that they follow robots.txt. Live fetch agents such as ChatGPT-User and Perplexity-User may ignore it, because a user triggered the request.

How do I recognize a fake GPTBot?

Check the IP address against OpenAI’s published file at openai.com/gptbot.json. A request with the GPTBot user agent from an address outside those ranges comes from someone else using the name.

Does a robots.txt Disallow keep a page out of Google?

A Disallow rule stops crawling of the content; the URL itself can still be indexed without a description if other pages link to it. To keep a page out of the index, use noindex and leave the page crawlable.

Sources

0%
Get insider knowledge first!
taismo Logo

© taismo GmbH

Address


Weißenfelder Str. 6
85551 Kirchheim near Munich, Germany