A crawler looks for groups with its own name in User-agent:, case does not matter, and combines them; with none, it uses the * group.
Of the Allow: and Disallow: lines that match the path, the longest wins. A tie goes to Allow. * matches any run of characters and $ the end of the address.
No file (404) means everything may be fetched; a server error (5xx) means crawlers treat everything as blocked. /robots.txt itself is always allowed.
AI crawlers, by what they do.
Training: GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Gemini), CCBot (Common Crawl), Applebot-Extended. Blocking them asks that your pages not be used to train models.
Search: OAI-SearchBot, Claude-SearchBot, PerplexityBot. Blocking them can keep your site out of those assistants' answers.
On a user's request: ChatGPT-User and Claude-User fetch a page someone asked about; OpenAI notes robots.txt may not apply to those.
Blocked is not hidden.
robots.txt stops crawling, not indexing: a blocked page can still show in search results, without a description, if other pages link to it. To keep a page out of results, let it be crawled and add noindex. Never rely on robots.txt for anything private; the file itself is public.
Questions people ask.
How does robots.txt decide whether a page is blocked?
A crawler uses the group for its own name, or the * group if it has none. Of the Allow and Disallow lines that match the path, the longest one wins; when an Allow and a Disallow are equally long, Allow wins (RFC 9309).
How do I block AI crawlers but keep Google?
Give each AI crawler its own group with Disallow: /, for example User-agent: GPTBot, User-agent: ClaudeBot, User-agent: Google-Extended, User-agent: CCBot. Googlebot keeps following its own rules, so search is unaffected.
Does Disallow keep a page out of Google?
No. It stops crawling, but a blocked page can still appear in results if other sites link to it. To keep a page out, allow crawling and add a noindex meta tag or header.
What happens if robots.txt is missing or broken?
A missing file (404) means crawlers may fetch everything. A server error (5xx) makes crawlers treat the whole site as blocked until it is fixed.
Is there a size limit?
Crawlers must read at least 500 KiB, and Google stops there: rules further down are ignored.
Hivex index
Short names, still free to register.
Starting something new? Hivex keeps a live index of short, brandable .si names nobody has claimed yet, each checked with the registry.
Hivex's free JSON API and MCP server check domains, DNS and registration records from your own code or from AI assistants that speak MCP. No key needed.
By Hivex. Updated 10 October 2026. Addresses are fetched live from Hivex's servers; page contents are never read or stored.
http://bbc.com/robots.txt
Checked 06:49:40 UTC · Fetched from Hivex's servers
Tested path
/
Groups
40
Sitemaps
39
Size
6.0 KiB
Note: 8 of 10 AI crawlers are blocked from /: Google-Extended, GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, CCBot, Applebot-Extended.
Good: Google may crawl /.
Who may fetch /
RFC 9309 rules: longest match wins
Crawler
/
Decided by
Search engines
GooglebotGoogle Search
Allowed
no rule matches (group *)
BingbotBing (and search built on it)
Allowed
no rule matches (group *)
AI training
Google-ExtendedGoogle: Gemini training and grounding
Blocked
line 164: Disallow: /
GPTBotOpenAI: model training
To test another page, put its path in the address you check, such as example.com/private/page.
3# The BBC's Terms of Use: https://www.bbc.co.uk/terms
4# - Explain the rules for using our services
5# - Tell you what you can do with our content
6#
7# In short: Please use our site like a human, not a robot.
8# That means:
9# - No scraping, crawling, or systematic extraction of content
10
Blocked
line 154: Disallow: /
ClaudeBotAnthropic: model training
Blocked
line 133: Disallow: /
CCBotCommon Crawl: open dataset many models train on
Blocked
line 121: Disallow: /
Applebot-ExtendedApple: AI training
Blocked
line 151: Disallow: /
AI search
OAI-SearchBotOpenAI: ChatGPT search
Blocked
line 169: Disallow: /
Claude-SearchBotAnthropic: Claude's search
Allowed
no rule matches (group *)
PerplexityBotPerplexity: search index
Blocked
line 174: Disallow: /
AI, fetched for a user
ChatGPT-UserOpenAI: pages a ChatGPT user asks for
Blocked
line 159: Disallow: /
Claude-UserAnthropic: pages a Claude user asks for
Allowed
no rule matches (group *)
# - No use of BBC content for training or fine-tuning AI models, including large language models (LLMs)
11# - No retrieval-augmented generation (RAG), AI-powered search, agentic AI or grounding using BBC content
12# - No creating datasets from BBC content
13# - No text and data mining (TDM) under Article 4 of the EU Directive on Copyright in the Digital Single Market
14# - No using BBC content to create summaries for your own use
15# - No business use without permission (details: https://www.bbc.co.uk/usingthebbc/terms/can-i-use-bbc-content-for-my-business/)
16# - The BBC reserves all rights in its content and expressly opts out of any statutory exceptions in any jurisdiction for text and data mining, as permitted by law
17
18# TL;DR: Browse, read, watch, enjoy - like a human.