What OpenAI's crawlers mean for SEO

OpenAI runs more than one bot. GPTBot is training, OAI-SearchBot is ChatGPT search, ChatGPT-User is a live fetch. What robots.txt can and cannot do.

Silkra team5 min read

Summarize

OpenAI does not have one crawler. It publishes several user agents, and they do different jobs. Treating "the GPTBot" as a single switch for "ChatGPT" is how a site gets removed from search answers while trying to opt out of training, or the reverse.

The source of truth is OpenAI's crawler overview. What follows is that page, in SEO language, plus the mistakes we keep seeing. It is not a leak. The products will move. They have an RSS feed on the docs for a reason.

There are four agents, not one

As of the docs we are reading, OpenAI lists four user agents.

OAI-SearchBot is for search. It is used to surface websites in search results in ChatGPT's search features. OpenAI says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links.

GPTBot crawls content that may be used in training generative AI foundation models. Disallowing GPTBot is how a site asks that this content stay out of that training pile.

ChatGPT-User is not an automatic crawl. When a person asks ChatGPT or a Custom GPT a question, the product may visit a page with this agent. It is also used when people hit GPT Actions. OpenAI says it is not used to decide whether content appears in Search.

OAI-AdsBot visits landing pages submitted as ads on ChatGPT, to check policy and relevance. OpenAI says that data is not used to train foundation models, and that the bot only visits pages submitted as ads.

The settings are independent. OpenAI's own example: allow OAI-SearchBot so the site can appear in search results, and disallow GPTBot so the crawl is not used for training. If both automated bots are allowed, they may reuse one crawl for both jobs so they do not fetch twice.

A robots.txt change for search can take about 24 hours to show up in their systems.

Notebook sketch of three small stamps labeled search, train, and fetch. The search stamp is the steel one.

Search, train, fetch. Same company. Three doors.

What each one means in an audit

For most SEO work, the first question is search, not training.

If the client wants to be eligible for ChatGPT search answers, OAI-SearchBot needs to be allowed, and the published IP ranges need to reach the site. OpenAI says so on the docs page. They also publish searchbot.json for those IPs.

If the client wants to opt out of foundation-model training, that is GPTBot and gptbot.json. It is a policy choice. It is not a ranking lever, and it is not the search opt-out.

If logs show ChatGPT-User, someone asked the product to look at a page, or an Action hit an endpoint. OpenAI says robots.txt rules may not apply because the fetch is user-initiated. Blocking it in robots.txt is not a reliable search control. Search opt-out is still OAI-SearchBot.

OAI-AdsBot only matters if someone is submitting ads. Most technical audits can ignore it until that is true.

User-agent strings include a version (OAI-SearchBot/1.4, GPTBot/1.4) and, for SearchBot, a Chrome-looking prefix. OpenAI says the version number may change. Match on the token, not the whole string. When they fetch robots.txt itself they may add a robots.txt marker in the UA so log files that omit paths are still readable.

What robots.txt can and cannot do

It can split training from search. That is the whole point of two tags.

It cannot stop a person from pasting a URL into ChatGPT. ChatGPT-User may still visit. OpenAI is explicit that this agent is not the Search switch.

It cannot make a page appear in answers. Allowing OAI-SearchBot is eligibility for their search crawl, not a ranking guarantee. The fetch still has to return the article.

It cannot rewrite a User-agent: * mistake you have not looked at. A blanket Disallow: / that was meant for a staging leftover will apply to these tokens unless a more specific group says otherwise. Read the file. Do not assume last year's "block the AI bots" snippet did what the slide promised.

A teaching pattern, copied from their independence example, not from a vendor pack:

code

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

That is "please do not train on this, please do crawl it for search." Whether you want that mix is a client conversation. The file is how you record the answer.

How this shows up in logs

The useful check is not a vibe. It is: which token, which IP, which path.

OpenAI publishes IP lists for each agent. Confirming a hit against searchbot.json, gptbot.json, chatgpt-user.json, or adsbot.json is how you avoid treating a spoofed UA as the real thing.

SearchBot's UA can look like a Mac Chrome session with a compatible; OAI-SearchBot/… tail. A naive "starts with Mozilla" filter will file it under humans. We have seen that misread in log reviews. Look for the token.

Then ask the technical questions. Did that path 200? Did the extract contain the article? Was the host production? Trusting the crawl still applies when the crawler is someone else's.

Silkra will show you whether a site blocks AI bots in robots.txt. Read that issue against this split. A GPTBot deny is not automatically "we are invisible in ChatGPT search." An OAI-SearchBot deny is the one that matches OpenAI's search opt-out language.

The part that is still fuzzy

We are trusting their labels. "This crawl is for search" and "this crawl is for training" are claims about an internal pipeline. The docs are the best public description. They are not a contract you can inspect.

ChatGPT-User is the foggiest in practice. User-initiated is clear as a category. How often it fires, what it stores, and how that text is used in the next turn are not on the page.

For SEO, the durable move is still small. Know which token does which job. Do not use one Disallow as a personality. Allow the search crawl if the client wants to be eligible. Fix the fetch so the article is what arrives. The rest is what happens after the lookup, and we are still sketching that.

Put crawl evidence to work

Download Silkra and turn audits, briefs, and fixes into one focused workflow.

Get started

Create your first workspace.

Crawl a site, then ask what needs attention.

Free to start. No credit card needed.