Asosiy kontentga o'tish
BiznesMarketing
Sign in
MarketingAugust 15, 2026· 3 min read

OpenAI Says robots.txt May Not Apply to the ChatGPT Bot

According to the "State of the Bots" report from TollBit, which tracks AI bot activity, ChatGPT-User — OpenAI's bot that fetches pages on request — is blocked by websites more often than comparable bots. Yet it's also the bot that accesses blocked pages the most. OpenAI maintains that robots.txt rules (a special file specifying which parts of a website bots may access) may not apply to this bot, since the page access is triggered by a user request.

According to TollBit's report for the first half of 2026, roughly 15% of AI bots detected on European sites accessed addresses marked as disallowed. This is concentrated in a handful of specific bots: ChatGPT-User, Bytespider, and Youbot each accessed disallowed pages on nearly half of the European sites that had explicitly blocked them. Among these, ChatGPT-User was the bot that accessed the most sites overall.

Newer bots, by contrast, are barely blocked at all. Claude-User is blocked by just 9% of European sites, versus 26% in North America. For Perplexity-User, those figures are 13% and 26% respectively. In Europe, most of the newest bots see single-digit block rates, but ChatGPT-User stands out sharply from that pattern.

OpenAI's bot documentation states that ChatGPT-User visits a page when a user asks ChatGPT a question, and because that action is user-initiated, robots.txt rules may not apply to it. Perplexity applies similar logic, saying Perplexity-User typically disregards the file. Anthropic takes a different approach — the company has previously said all three of its bots comply with robots.txt. TollBit, for its part, counts any request to a disallowed address as a violation, regardless of the bot operator's position.

A key nuance here is that, per OpenAI's own documentation, whether a site appears in ChatGPT search results is determined by a different bot — OAI-SearchBot — not ChatGPT-User. That means sites blocking both bots to protect against AI traffic are effectively giving up search visibility while only retaining control over page fetches under a special exception.

A site's server logs or CDN (content delivery network) records show what actually happened, while the robots.txt file only reflects what was requested. Cloudflare is updating its own bot-control system, shifting the decision to the network level. For recognized bots, compliance will no longer depend on the bot itself. Starting September 15, new domains joining Cloudflare will have Training- and Agent-type bots blocked by default on pages carrying ads, while Search-type bots will continue to be allowed.

Whether this user-initiated exception holds up in the long run remains an open question. It comes down to the distinction between a user requesting a page and a bot crawling it on its own — and right now, all the major AI assistants are fetching pages precisely the user-initiated way.

Source: Search Engine Journal · view original article
Ulashish:TelegramLinkedIn
← Back to homepage

Related articles