15% of AI page fetchers ignored European robots.txt blocks, TollBit reports

15% of AI page fetchers in Europe reached disallowed URLs, TollBit finds

Roughly one in seven AI page-fetching agents crawling European websites ignored explicit instructions to stay away, according to new measurement data published August 14, 2026. The most frequent offender belongs to ChatGPT, the world’s most widely used chat assistant.

The finding appears in TollBit’s latest State of the Bots report covering the first half of 2026. Search Engine Journal’s Matt G. Southern reported the numbers, comparing them against crawler documentation published by the companies operating these bots.

The measurement focused narrowly on page fetchers, the agents that load a single URL in real time when someone asks a chat assistant a question requiring that page. Training crawlers and search indexers fell outside the scope.

Across European sites examined, approximately 15% of identified fetchers reached URLs marked as disallowed. The bypasses concentrated heavily among three agents. ChatGPT-User, Bytespider, and Youbot each accessed restricted pages on nearly half the European sites that had named them in robots.txt files.

ChatGPT-User holds an unenviable distinction. More sites disallow it than any other fetching bot, yet it also reaches disallowed pages on more sites than any competitor. The paradox stems from a simple reality: robots.txt carries no enforcement mechanism.

TollBit’s methodology counted any request to a disallowed URL as a bypass, regardless of what operator documentation claimed about compliance.

Meanwhile, European site owners block newer agents far less frequently than their North American counterparts. Only 9% of European sites disallow Claude-User, versus 26% in North America. Perplexity-User sits at 13% in Europe against 26% across the Atlantic. Most newest fetchers register single-digit disallow rates.

ChatGPT-User stands as the exception, blocked far more often, which explains its higher observed bypass count. An agent never named in robots.txt cannot violate a rule that was never written.

OpenAI’s developer documentation states plainly that ChatGPT-User may ignore robots.txt because a person initiated the request. Perplexity takes the same position. Anthropic maintains all three of its bots respect the file. Google formalised this category in March 2026 with Google-Agent, documenting that such fetchers generally bypass robots.txt.

Cloudflare’s July 1 policy update shifts enforcement to the network layer. Starting September 15, 2026, new domains onboarding to Cloudflare will block Training and Agent crawlers by default on ad-carrying pages. Search remains allowed.

The shift reflects a measurable gap between what robots.txt requests and what server logs record. That gap carries commercial weight as automated traffic increasingly distorts publisher metrics and advertiser campaign data.