How to add GPTBot to robots.txt
GPTBot and OAI-SearchBot are separate robots.txt controls with different purposes. GPTBot governs whether OpenAI may use your content for model training. OAI-SearchBot governs whether your pages can appear in ChatGPT search results. If you want ChatGPT search visibility, allow OAI-SearchBot. Allowing GPTBot is a training-data decision, not a search-visibility fix.
Exact robots.txt configuration for OpenAI crawlers
OpenAI uses multiple distinct user-agents depending on the context of the retrieval. GPTBot crawls web pages to build training datasets, while OAI-SearchBot powers real-time web search and grounding for ChatGPT users.
Choose each policy at the root of your robots.txt file based on its documented purpose. Search access and training permission are separate decisions:
| User-Agent | Purpose | Recommended Policy |
|---|---|---|
| OAI-SearchBot | Real-time ChatGPT Web Search | Allow: / |
| ChatGPT-User | Direct user-requested retrieval | Review separately |
| GPTBot | Model-training use | Owner policy choice |
Step-by-step implementation guide
Ensure your robots.txt file is served with an HTTP 200 status code and Content-Type: text/plain. Avoid using JavaScript-based redirects or dynamic authentication on the robots.txt endpoint.
- 1Open your public robots.txt file in your web root or Next.js app directory.
- 2Add 'User-agent: OAI-SearchBot' followed by 'Allow: /'.
- 3Do not allow GPTBot merely to pursue search visibility; it is a separate training-use policy.
- 4Disallow private application endpoints such as /api/, /dashboard/, and /auth/.
- 5Test your live robots.txt using HubSEO's free AI Crawler Access Checker.
Frequently Asked Questions
Does allowing GPTBot mean my content is scraped for AI training?
GPTBot and OAI-SearchBot are separate documented tokens. GPTBot concerns model-training use; OAI-SearchBot concerns ChatGPT search discovery. Allowing one does not prove access for the other.
Where should robots.txt be located?
Robots.txt must always reside at the apex root of your domain (e.g. https://yourdomain.com/robots.txt) to be parsed by crawlers.
How quickly does ChatGPT recognize robots.txt changes?
OpenAI crawlers typically refresh their cached robots.txt directives within 24 to 48 hours of publication.
Can I block specific subdirectories from GPTBot?
Yes. You can add specific 'Disallow: /private-folder/' lines beneath the 'User-agent: GPTBot' directive.