GEOTIMES
A record of who generative AI search recommends
Home / International
International

Publishers Block Training Bots, Allow Answering Bots

Publishers are blocking training crawlers like GPTBot while leaving search-answer bots open, new data shows.

Source TechnologyChecker.io · 2026-09-03 · Checked 2026-09-18

Publishers Block Training Bots, Allow Answering Bots

A study pulled robots.txt files from a sample of top domains on Cloudflare's network to measure block-and-allow ratios by AI bot. According to TechnologyChecker, as of August 31, 2026, GPTBot's disallow-to-allow ratio stood at 696 to 299 entries, or 2.33 to 1. ClaudeBot came in at 619 to 259, or 2.39 to 1. Training-focused crawlers were being blocked far more often than they were allowed.

OAI-SearchBot, which is used for search answers, showed the opposite pattern at 250 to 266, or 0.94 to 1, meaning it was allowed more than it was blocked. Googlebot was close to even at 307 to 277, or 1.11 to 1.

The source summed this up as publishers not blocking "AI" but blocking training while allowing answering, and doing so even to the same company. The basis for that split is how much traffic a bot's crawling actually turns into real visits.

In the first quarter of 2026, OpenAI's GPTBot pulled 1,255 pages for every visit it generated. The company's own OAI-SearchBot, by contrast, kept generating real visits in proportion to what it crawled.

By the source's own standard, a ratio above 2.0 to 1 marks a training or training-opt-out bot, while anything below 1.2 to 1 marks a search or answer bot. This also shifted over time. Anthropic crawled 20,583 pages per visit in Q1 2026, improved to 1,917 to 1 by July, and held roughly steady through the end of August.

OpenAI's GPTBot improved similarly, from 1,255 to 1 in Q1 to 251 to 1 by July. Meta-ExternalAgent, on the other hand, still generated zero visits. Alongside this, the share of AI crawler traffic that was training-purpose fell from 89.4% in Q1 2026 to 40.15% in August.

By industry, retail accounted for 28.1% of AI crawling traffic, the largest share, but blocked selectively. News and media made up just 2.7% of that traffic yet blocked the most aggressively. Among tech-sector domains, 910 had added AI-specific block rules to their robots.txt.

In Korea

This study sampled top domains on Cloudflare's network, and the source doesn't break out numbers for Korean sites specifically. Publishers or online retailers here can only find out how their own robots.txt treats training bots like GPTBot and ClaudeBot versus answer bots like OAI-SearchBot and PerplexityBot by opening the file and checking directly.

The source advises publishers to block bots with a crawl-to-visit ratio above 1,000 to 1 and allow those below 200 to 1, noting that robots.txt is a voluntary standard bots choose to follow, so actually enforcing access limits requires firewall rules as well.

Source: TechnologyChecker.io, updated September 3, 2026 (data as of August 31, 2026). Checked September 18, 2026. Read the original