---
title: What Is GPTBot? How AI Crawlers Read Your Site
description: GPTBot is OpenAI's web crawler. Learn what it reads, how it differs from other AI crawlers, and how robots.txt lets you block GPTBot or allow it.
image: https://www.finemediabw.com/hubfs/Blog/AEO/Awareness/What%20Is%20GPTBot%20-%20How%20AI%20Crawlers%20Read%20Your%20Site/gptbot-share-1200x630.png
---

<https://www.finemediabw.com/blog/gptbot#body>

[![Fine\_Media\_logo\_Set\_HubSpot Website Logo Stacked](https://www.finemediabw.com/hubfs/Fine_Media_logo_Set_HubSpot%20Website%20Logo%20Stacked.svg "Fine_Media_logo_Set_HubSpot Website Logo Stacked")](https://www.finemediabw.com/)

- Solutions 
  
    - [Inbound Marketing](https://www.finemediabw.com/inbound-marketing)
    - [Website Design](https://www.finemediabw.com/website-design-and-development)
    - [Sales Enablement](https://www.finemediabw.com/sales-enablement)
    - [Customer Success](https://www.finemediabw.com/customer-success)
    - [HubSpot Services](https://www.finemediabw.com/hubspot-services)
- [Insights](https://www.finemediabw.com/blog)
- [About](https://www.finemediabw.com/about-us)
  
   Show submenu for About 
  
    - [About us](https://www.finemediabw.com/about-us)
    - [Meet the Team](https://www.finemediabw.com/our-team)
    - [Our Story](https://www.finemediabw.com/our-story)

- Solutions 
  
    - [Inbound Marketing](https://www.finemediabw.com/inbound-marketing)
    - [Website Design](https://www.finemediabw.com/website-design-and-development)
    - [Sales Enablement](https://www.finemediabw.com/sales-enablement)
    - [Customer Success](https://www.finemediabw.com/customer-success)
    - [HubSpot Services](https://www.finemediabw.com/hubspot-services)
- [Insights](https://www.finemediabw.com/blog)
- [About](https://www.finemediabw.com/about-us)
  
   Show submenu for About 
  
    - [About us](https://www.finemediabw.com/about-us)
    - [Meet the Team](https://www.finemediabw.com/our-team)
    - [Our Story](https://www.finemediabw.com/our-story)

- [Contact us](https://www.finemediabw.com/contact)

- [Contact us](https://www.finemediabw.com/contact)

![What Is GPTBot? How AI Crawlers Read Your Site. Dark and slate type on a white ground, Propello.](https://www.finemediabw.com/hubfs/Blog/AEO/Awareness/What%20Is%20GPTBot%20-%20How%20AI%20Crawlers%20Read%20Your%20Site/gptbot-share-1200x630.png)

 Oct 10, 2026, 1:36:08 PM | [AEO & AI Search](https://www.finemediabw.com/blog/tag/aeo-ai-search)

# What Is GPTBot? How AI Crawlers Read Your Site

GPTBot is OpenAI's web crawler. Learn what it reads, how it differs from other AI crawlers, and how robots.txt lets you block GPTBot or allow it.

Share

- [mailto:?&subject=What%20Is%20GPTBot?%20How%20AI%20Crawlers%20Read%20Your%20Site&body=What%20Is%20GPTBot?%20How%20AI%20Crawlers%20Read%20Your%20Site%0A(https%3A%2F%2Fwww.finemediabw.com%2Fblog%2Fgptbot)](mailto:?&subject=What%20Is%20GPTBot?%20How%20AI%20Crawlers%20Read%20Your%20Site&body=What%20Is%20GPTBot?%20How%20AI%20Crawlers%20Read%20Your%20Site%0A(https%3A%2F%2Fwww.finemediabw.com%2Fblog%2Fgptbot))
- <https://www.linkedin.com/shareArticle?mini=true&url=https%3A%2F%2Fwww.finemediabw.com%2Fblog%2Fgptbot&title=What%20Is%20GPTBot?%20How%20AI%20Crawlers%20Read%20Your%20Site&summary=&source=>
- <https://twitter.com/intent/tweet?text=What+Is+GPTBot%3F+How+AI+Crawlers+Read+Your+Site&url=(https%3A%2F%2Fwww.finemediabw.com%2Fblog%2Fgptbot)>
- <https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fwww.finemediabw.com%2Fblog%2Fgptbot>

If you run revenue at a B2B company, you may have seen GPTBot in your server logs or heard you should block it. GPTBot is OpenAI's web crawler: it reads publicly accessible pages so OpenAI can train its large language models and other AI models, and website owners control it with robots.txt.

This article is for the CEOs, founders, CROs and revenue leaders who own the answer. You will see how GPTBot works, how it differs from OpenAI's other bots, which other AI crawlers visit your site, and how to choose what to allow. It builds on [our guide to AEO](https://www.finemediabw.com/blog/ultimate-guide-ai-engine-optimization).

 

| **GPTBot** GPTBot is the web crawler OpenAI uses to collect publicly accessible content that may be used to train its generative AI foundation models, and site owners control it through robots.txt. |
| --- |

## How GPTBot reads your site

OpenAI says GPTBot exists to make its generative AI foundation models more useful and safe. It behaves like other web crawlers: it requests your web pages over the open web, reads publicly accessible content and publicly available data that any visitor without a login could read, and moves from link to link.

It does not log in, so it does not see anything you keep behind a login or a paywall. It also does not index pages for search results; that job belongs to a different OpenAI bot, covered below.

Every request carries a user agent string, a line of text that names the software making the request. OpenAI's [crawler documentation](https://developers.openai.com/api/docs/bots) gives an example for GPTBot that includes the user agent token GPTBot and a version number, and says the version may change.

OpenAI lists the full user agent string, so your team can match it in your logs and see every one of the GPTBot visits.

Before it reads a page, a well-behaved crawler first fetches your robots.txt file, a plain txt file at the root of your site. GPTBot looks for the section addressed to its own token and follows the rules there.

That is the whole control system, and it is why one small file decides what an AI bot may read.

OpenAI also publishes the IP address ranges its bots use, so your technical team can confirm a visit really came from OpenAI and not from software pretending to be GPTBot. Anyone can copy a user agent string, so that check matters.

## The AI crawlers you will see in your logs

![Your site at the centre with six crawler roles around it: training crawlers, a Google training token, ChatGPT search, other AI search bots, user fetchers and Google Search.](https://www.finemediabw.com/hs-fs/hubfs/Blog/AEO/Awareness/What%20Is%20GPTBot%20-%20How%20AI%20Crawlers%20Read%20Your%20Site/gptbot-ai-crawler-roles-blog-1600x900.png?width=1600&height=900&name=gptbot-ai-crawler-roles-blog-1600x900.png)

GPTBot is one of several AI bots, and crawlers from other AI systems visit too. They do not do the same job, and the operators do not treat them the same way.

A rule written for one does not govern the others. The jobs fall into three groups.

| Crawler | Operator | What it is for | Does robots.txt control it? |
| --- | --- | --- | --- |
| GPTBot | OpenAI | Collecting content that may train generative AI models | Yes |
| OAI-SearchBot | OpenAI | Surfacing sites in ChatGPT search | Yes |
| ChatGPT-User | OpenAI | Visiting a page when a user asks | Not reliably |
| ClaudeBot | Anthropic | Collecting content that may contribute to training | Yes |
| Claude-SearchBot | Anthropic | Improving the relevance of search answers | Yes |
| Claude-User | Anthropic | Fetching pages for user requests | Yes, per Anthropic |
| PerplexityBot | Perplexity | Surfacing and linking sites in Perplexity search | Yes |
| Perplexity-User | Perplexity | Loading a page when a user asks | Generally ignores it |
| Googlebot | Google | Crawling for Google Search | Yes |
| Google-Extended | Google | A token that limits AI training and grounding use | Yes |

### Crawlers that gather data to train AI models

GPTBot and ClaudeBot belong here. They gather data from your pages, and the data collected may become training data for large language models, the machine learning systems behind conversational AI, including future AI models.

Anthropic describes ClaudeBot as collecting web content that could potentially contribute to training. It says restricting it signals that your future materials should be left out of training datasets.

Google-Extended is a token, not a separate crawler. Google's [crawler documentation](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers) describes it as a standalone product token publishers can use to manage whether content Google crawls is used for AI training and grounding in Gemini Apps and the Vertex AI API for Gemini.

It does not affect inclusion in Google Search or work as a ranking signal. Most search engine bots, including Googlebot, follow robots.txt for automatic crawls.

### Crawlers that power AI search

OAI-SearchBot, Claude-SearchBot and PerplexityBot index pages so an AI engine can show or cite them in answers. Perplexity's [bot documentation](https://docs.perplexity.ai/guides/bots) says PerplexityBot surfaces and links websites in its search results and is not used to gather content for AI foundation models.

Blocking a search bot can keep you out of that engine's AI generated responses. That matters when users search with an assistant instead of a list of links.

### Fetchers that act for a user

ChatGPT-User, Claude-User and Perplexity-User load a page because a person asked a question. They are not crawling on a schedule.

OpenAI says robots.txt rules may not apply to user-initiated actions, and Perplexity says its user fetcher generally ignores robots.txt. Plan for that when you decide what a robots.txt file can promise.

## GPTBot is not the same as OAI-SearchBot or ChatGPT-User

Many site owners treat GPTBot as the only OpenAI bot. It is not, and the difference changes what a block does.

OpenAI says each setting is independent of the others. A site can allow OAI-SearchBot to appear in search results while disallowing GPTBot to signal that its content should not be used for training.

In plain terms: disallowing GPTBot is a statement about training. It is not a statement about whether ChatGPT may show your pages in answers.

That second question belongs to OAI-SearchBot. OpenAI says sites that opt out of it will not be shown in ChatGPT search answers, though they can still appear as navigational links.

For search changes, OpenAI notes it can take about 24 hours after you update robots.txt for its systems to adjust.

## How to block GPTBot and other AI bots with robots.txt

![A table of five crawlers, GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and Google-Extended, with what happens if you allow each, what happens if you block it and what that does.](https://www.finemediabw.com/hs-fs/hubfs/Blog/AEO/Awareness/What%20Is%20GPTBot%20-%20How%20AI%20Crawlers%20Read%20Your%20Site/gptbot-robots-txt-decision-path-blog-1600x900.png?width=1600&height=900&name=gptbot-robots-txt-decision-path-blog-1600x900.png)

A robots.txt file is a set of requests, written per user agent. Each block starts with the user agent token of the crawler it addresses and then lists paths it may or may not read.

You can block GPTBot specifically without touching other crawlers. Any rule to block GPTBot only affects GPTBot, so every one of the other crawlers you care about needs its own lines. A robots.txt rule tells web crawlers what they may read from now on.

To keep your entire site out of OpenAI's training crawl while staying visible in ChatGPT search, a file could read like this:

```
User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /
```

You can also protect part of a site. A rule such as \`Disallow: /pricing-research/\` under the GPTBot token blocks only that folder.

Anthropic's documentation tells you to add the rule for each subdomain. It also says it supports the non-standard Crawl-delay setting if you want to slow a crawler rather than stop it.

Two limits are worth knowing. First, robots.txt is a request that reputable operators honor, not a lock. Anything truly confidential belongs behind a login.

Second, Anthropic advises against blocking by IP address as an opt-out method, saying it may not work reliably. An IP block also means keeping an updated list of OpenAI's published addresses.

If a crawler's load becomes a problem, your technical team can rate limit its requests. It can also return a 403 Forbidden response at the server, for example with a rule in the .htaccess file on an Apache web server.

## Decide whether to block GPTBot or allow GPTBot access

The right setting depends on why your site exists, what you sell and how potential customers research. If you prioritize content control, say so; if you prioritize reach, say that instead. Either way, treat it as a business decision about visibility, not a technical default.

Some publishers also raise ethical concerns and an intellectual property angle about training use, and those are fair reasons to block. A few questions sort most cases.

### Why some site owners block GPTBot and others allow it

Site owners who block GPTBot usually give the same reasons: control over proprietary content, an intellectual property concern, or a view that training on their pages should be their decision.

Large publishers do block it. As of 10 October 2026, the New York Times' own [robots.txt file](https://www.nytimes.com/robots.txt) disallowed GPTBot for the whole site, and The New York Times is one of many news outlets that do.

A study of the top one million most-visited websites, by Paul Bouchaud and Pedro Ramaciotti, found that a quarter of the top thousand sites restrict AI crawlers, falling to one-tenth across the million. It also found that 34.2% of news outlets disallow GPTBot, in [an October 2025 arXiv paper](https://arxiv.org/abs/2510.09031).

Some marketers worry that a block could undermine SEO efforts, but a rule for GPTBot does not touch Google Search.

The legal side is unsettled. GPTBot operates in a gray area of copyright law, and arguments about AI training and fair use are still being tested.

Rules such as GDPR and CCPA may matter if your public pages hold personal data. Ask your legal team before you decide, because a crawler policy is also a risk policy.

Site owners who allow GPTBot give a different reason. They want AI models to know their category, their products and their wording, so that AI generated responses describe them correctly.

Allowing GPTBot to crawl costs nothing to try, and you can block GPTBot later by editing one file. Allowing GPTBot also supports your generative engine optimization work, because the data collected can help models represent your brand accurately. Neither choice is permanent.

### Is your content a product or a way to be found?

If your site hosts user generated content, such as community posts, remember that GPTBot reads whatever is publicly accessible, whoever wrote it.

If you publish proprietary content that customers pay for, such as data, research or training, you have a clear reason to block training crawlers and keep that material behind access controls.

If your articles exist to be found by buyers, then allowing GPTBot to crawl, with GPTBot access limited to the sections you choose, may help AI systems describe your category and your company accurately. Letting GPTBot crawl a public blog is a different decision from letting GPTBot crawl a gated library.

### Which engines should be able to cite you?

Blocking a training crawler and blocking a search crawler are different choices. Blocking access to everything is rarely what a B2B company wants.

A company that wants AI visibility usually wants its pages open to search bots even if it limits training. Allowing GPTBot is a separate choice from allowing search bots, and letting GPTBot crawl says nothing about whether a search engine can show you.

Our piece on [AI search ranking factors](https://www.finemediabw.com/blog/ai-search-ranking-factors) explains how engines pick sources to cite when they answer user questions about your category.

### Does it change how you appear in Google?

Not through GPTBot. Google's documentation on its AI features says it offers no special controls for Google's AI Overviews and AI Mode.

You manage them with the same Search controls as the rest of Google Search, such as Googlebot rules and the nosnippet and noindex directives. The page ties Google-Extended to training and grounding in other systems.

## How to check what is visiting your site

Any automated system accessing your site leaves a trace. Start with your server logs or your bot management tool, and filter for GPTBot and the other tokens in the table above.

Google Search Console reports on Google's own crawling, so it will not show these visits. Add the check to your regular site monitoring, and watch site performance if a crawler's requests ever spike.

A visit with a matching user agent string tells you a crawler named itself that way. Confirm it against the operator's published IP ranges before you trust it.

Then read your own robots.txt file as a stranger would. List every AI crawler it mentions, every one it ignores, and any blanket rule that blocks all bots by accident.

Finally, write the decision down with a date and an owner. Operators add and rename bots, and the documentation pages are the source of truth, so review the file when they change.

## What you gain when this is done properly

![A crawl rule at the centre with three arcs for what sales, marketing and customer success each gain from deciding on purpose what AI crawlers may read.](https://www.finemediabw.com/hs-fs/hubfs/Blog/AEO/Awareness/What%20Is%20GPTBot%20-%20How%20AI%20Crawlers%20Read%20Your%20Site/gptbot-what-each-team-gains-blog-1600x900.png?width=1600&height=900&name=gptbot-what-each-team-gains-blog-1600x900.png)

Treating crawler access as a deliberate policy helps each team that works with the pages you publish.

### Sales

Some prospects ask AI assistants about your category before they speak to you. When the pages you want read are open to search bots, those answers can point to your own material instead of a competitor's.

Reps then spend less of a first call correcting basics.

### Marketing

Marketing gets a clear rule for which content is meant to be found and which is protected. That ends ad hoc blocking decisions and protects premium assets.

It also keeps traditional search and AI search visibility working from the same plan.

### Customer success

Customers ask AI tools how to use your product, and AI tools answer from whatever pages they can reach. If your help content is open to the right crawlers, those tools can draw on your current instructions instead of outdated or inaccurate sources.

Your team then fields fewer questions caused by a wrong answer.

## Decide on purpose what AI crawlers may read

GPTBot is a training crawler with a single control: a rule in robots.txt. You can allow GPTBot, block GPTBot or give GPTBot access to part of your site.

The larger picture is a wider AI ecosystem of bots with different jobs, and the choice is which of those jobs you want your site to take part in.

Propello designs and builds connected GTM systems on HubSpot. If you want your content and AI visibility planned as part of one revenue system, an audit is the place to start.

[Book a Propello GTM Audit](https://www.finemediabw.com/contact)

## Frequently asked questions

 What is GPTBot?

GPTBot is OpenAI's web crawler. It reads publicly accessible pages to collect content that may be used to train OpenAI's generative AI foundation models. You identify it by the GPTBot token in the user agent string, and you control it with a rule in your robots.txt file.

 Should I block GPTBot?

It depends on your content. Block it if your pages are proprietary products or you do not want them used for training. Allow it if your pages exist to be found and you want AI models to describe your category accurately. Blocking GPTBot does not remove you from ChatGPT search.

 How do I block GPTBot?

Add a section to robots.txt that names the GPTBot user agent and disallows the paths you want protected, using a slash to cover the whole site. Leave OAI-SearchBot allowed if you still want ChatGPT search to show you. OpenAI says search changes can take about 24 hours.

 Does blocking GPTBot hurt my Google rankings?

No. GPTBot belongs to OpenAI, not Google, and Google says its Google-Extended token does not affect inclusion in Google Search or act as a ranking signal. Google Search is managed through Googlebot rules and snippet and index controls, not through any OpenAI setting.

 Does robots.txt stop every AI crawler?

No. Reputable operators honor it for their automatic crawlers, but OpenAI says robots.txt rules may not apply to user-initiated actions, and Perplexity says its user fetcher generally ignores them. Treat robots.txt as a policy signal, and keep truly confidential content behind a login.

![Tumisang Bogwasi](https://app.hubspot.com/settings/avatar/77d7e2eaad8ff71b24463dcc39a31e9e)

### Written By: Tumisang Bogwasi

Tumisang Bogwasi is the founder and CEO of Propello, a HubSpot partner that designs and builds connected go-to-market systems.

[mailto:tumib@finemediabw.com](mailto:tumib@finemediabw.com) <https://www.linkedin.com/in/tumisangbogwasi>

## You May Also Like

[Learn More](https://www.example.com)

[![Fine\_Media\_logo\_Set\_HubSpot Website Logo Stacked White BG White-1](https://www.finemediabw.com/hubfs/Fine_Media_logo_Set_HubSpot%20Website%20Logo%20Stacked%20White%20BG%20White-1.svg "Fine_Media_logo_Set_HubSpot Website Logo Stacked White BG White-1")](https://www.finemediabw.com)

We help you propel growth at every stage of the customer journey with inbound marketing and Hubspot.

![provider-horizontal-white](https://www.finemediabw.com/hubfs/provider-horizontal-white.svg "provider-horizontal-white")

Blogs

[What is Inbound Marketing](https://www.finemediabw.com/blog/what-is-inbound-marketing)

[Why work with a HubSpot Agency](https://www.finemediabw.com/blog/why-work-with-hubspot-agency)

[Benefits of a Website](https://www.finemediabw.com/blog/the-power-of-website-design-a-cornerstone-of-brand-identity)

[Ultimate Guide to Closing more Deals](https://www.finemediabw.com/blog/how-to-close-more-deals-with-sales-enablement)

[Why Every Business Needs Customer Success](https://www.finemediabw.com/blog/why-customer-success-is-important)

Services

 

[Inbound Marketing](https://www.finemediabw.com/inbound-marketing)

[Website Design](https://www.finemediabw.com/website-design-and-development)

[Sales Enablement](https://www.finemediabw.com/sales-enablement)

[Customer Success](https://www.finemediabw.com/customer-success)

[HubSpot Services](https://www.finemediabw.com/hubspot-services)

Quick Links

 

[Home](https://www.finemediabw.com/)

Services

[About us](https://www.finemediabw.com/about-us)

[Contact](https://www.finemediabw.com/contact-us)

© 2026 Fine Media (Pty) Ltd. All rights reserved. [Legal Notice](https://www.finemediabw.com/user-agreement) | [Privacy Policy](https://www.finemediabw.com/privacy-policy) 

<https://facebook.com> <https://twitter.com> <https://www.instagram.com/finemediabw/> <https://linkedin.com> <https://www.pinterest.com/>

```json
{
  "@context" : "https://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Tumisang Bogwasi",
    "url" : "https://www.finemediabw.com/blog/author/tumisang-bogwasi"
  },
  "datePublished" : "2026-10-10T11:36:08.000Z",
  "headline" : "What Is GPTBot? How AI Crawlers Read Your Site",
  "image" : [ "https://www.finemediabw.com/hubfs/Blog/AEO/Awareness/What%20Is%20GPTBot%20-%20How%20AI%20Crawlers%20Read%20Your%20Site/gptbot-share-1200x630.png" ],
  "mainEntityOfPage" : {
    "@id" : "https://www.finemediabw.com/blog/gptbot",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "url" : "https://www.finemediabw.com/hubfs/Fine_Media_logo_Set%20copy_HubSpot_Logo_Long_Stacked.svg"
    },
    "name" : "Fine Media"
  }
}
```

```json
{
  "@context" : "https://schema.org",
  "@type" : "FAQPage",
  "mainEntity" : [ {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "GPTBot is OpenAI's web crawler. It reads publicly accessible pages to collect content that may be used to train OpenAI's generative AI foundation models. You identify it by the GPTBot token in the user agent string, and you control it with a rule in your robots.txt file."
    },
    "name" : "What is GPTBot?"
  }, {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "It depends on your content. Block it if your pages are proprietary products or you do not want them used for training. Allow it if your pages exist to be found and you want AI models to describe your category accurately. Blocking GPTBot does not remove you from ChatGPT search."
    },
    "name" : "Should I block GPTBot?"
  }, {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "Add a section to robots.txt that names the GPTBot user agent and disallows the paths you want protected, using a slash to cover the whole site. Leave OAI-SearchBot allowed if you still want ChatGPT search to show you. OpenAI says search changes can take about 24 hours."
    },
    "name" : "How do I block GPTBot?"
  }, {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "No. GPTBot belongs to OpenAI, not Google, and Google says its Google-Extended token does not affect inclusion in Google Search or act as a ranking signal. Google Search is managed through Googlebot rules and snippet and index controls, not through any OpenAI setting."
    },
    "name" : "Does blocking GPTBot hurt my Google rankings?"
  }, {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "No. Reputable operators honor it for their automatic crawlers, but OpenAI says robots.txt rules may not apply to user-initiated actions, and Perplexity says its user fetcher generally ignores them. Treat robots.txt as a policy signal, and keep truly confidential content behind a login."
    },
    "name" : "Does robots.txt stop every AI crawler?"
  } ]
}
```