What is GPTBot? AI crawler guide
GPTBot is a crawler or fetcher associated with OpenAI. See source-graded user-agent, robots.txt, verification, and observation guidance.
What is GPTBot?
GPTBot is a documented OpenAI crawler or fetcher. OpenAI crawler for content that may be used to improve generative AI foundation models.
Evidence status
| Field | Value |
|---|---|
| Evidence | Officially documented |
| Lifecycle | Active |
| Purpose | training |
| robots.txt posture | honors |
| Source checked | 2026-08-18 |
Documented user-agent
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
Allowing GPTBot creates the possibility of retrieval. To find out whether ChatGPT actually cites or recommends you, track your ChatGPT mentions over time.
To see why GPTBot access is worth checking rather than assuming, read what happened when one site fixed a page its crawlers could not read.
robots.txt allow example
User-agent: GPTBot Allow: /
robots.txt block example
User-agent: GPTBot Disallow: /
What the rule can and cannot do
Robots.txt expresses an access policy to compliant automated crawlers. It does not authenticate the sender, remove content already collected, or guarantee that a model will use or not use content.
How to verify a request
- Match the complete documented user-agent where one exists. A match is only a clue.
- Use the operator-published check: Published IP ranges.
- Keep the observed request, operator documentation, and any inference as separate fields.
Observed in Trakkr connected-site data
| Measure | Value |
|---|---|
| Matching signature | GPTBot |
| Classified requests | 254251 |
| Sites in sample | 85 |
| Window | 2026-07-18 through 2026-08-17 |
| Method | Finalized daily crawler summaries from connected Trakkr sites, grouped by the crawler signature detected in each request. |
| Limits | This is a connected-site sample, not a representative sample of the web. A matching user-agent or signature does not prove that the named operator sent the request. Counts describe classified request signatures, not market share or unique pages crawled. |
JavaScript behavior
The cited operator material does not verify JavaScript rendering. Serve useful HTML before client-side JavaScript where possible.
Related bots
- KimiBot: Also tracked as a training crawler.
- MistralAI-Training: Also tracked as a training crawler.
- Applebot-Extended: Also tracked as a training crawler.
- Meta-ExternalAgent: Also tracked as a training crawler.
- Bytespider: Also tracked as a training crawler.
- CCBot: Also tracked as a training crawler.
- Google-Extended: Also tracked as a training crawler.
- AI2Bot: Also tracked as a training crawler.
- ClaudeBot: Also tracked as a training crawler.
- GPTBot: GPTBot is the glossary definition behind this crawler guide.
- AI Training Opt-Out: GPTBot is a training crawler tied to this policy decision.
Frequently Asked Questions
What is GPTBot?
GPTBot is a officially documented crawler or fetcher record associated with OpenAI.
What user-agent does GPTBot use?
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot. Check the evidence status before attributing a matching request.
Can GPTBot be blocked in robots.txt?
The record identifies GPTBot as the token to review. Robots.txt is a request policy, not proof of model use or non-use.
How can I verify a GPTBot request?
Start with the full user-agent, then use published ip ranges from the operator. A user-agent match alone is not proof.
Data & Sources
- OpenAI documentation - Primary source for GPTBot crawler details.
- GPTBot verification: Published IP ranges - Verification information published for GPTBot.