Skip to main content

Every discussion about AI and your website ends up at the same file. Block GPTBot. Allow OAI-SearchBot. Add Google-Extended. The advice is everywhere, it is mostly correct, and it assumes something nobody checks: that a crawler asking your server for /robots.txt actually receives it.

On 6 October 2026 we asked 100 Irish websites for that file. Twenty-seven of them would not hand it over to a client that identified itself as a crawler. The list includes a government portal, three universities, a national health service, four banks and insurers, and most of the country’s large retailers.

TL;DR

  • We requested /robots.txt from 100 Irish domains. 76 served it, 22 returned a 4xx on both the apex and the www host, one timed out, and one no longer resolves.
  • RFC 9309 is explicit about what a 4xx means: “the crawler MAY access any resources on the server”. A 403 on your robots.txt is not protection, it is permission.
  • Five more sites served the file to a browser user agent and refused the documented OAI-SearchBot user agent, including gov.ie.
  • Of the 76 readable files, 57 name no AI agent at all, and the 19 that do are mostly blocking tokens the vendors no longer document: 10 block anthropic-ai, 8 block claude-web, 5 block perplexity-ai, and exactly one names the live Claude-SearchBot.
  • The directives that are live mostly point the wrong way: eight sites block the crawler that puts them in ChatGPT’s answers, while five block a user-triggered fetcher that OpenAI documents as “robots.txt rules may not apply”.

How we ran it

The sample is 100 Irish domains across news, banking, insurance, retail, state services, universities, transport, telecoms and classifieds. Each one got a single GET to https://www.<domain>/robots.txt with a current Chrome user agent string, following redirects. Anything that failed was retried on the apex host, and then retried again with the user agent string OpenAI publishes for OAI-SearchBot robots.txt fetches.

The results split cleanly. 76 domains returned 200 with a parseable file. 22 returned a 4xx on every attempt: 17 served 403, four served 404, and aerlingus.com answered a GET for robots.txt with 405 Method Not Allowed and a 12 KB body. One host, www.welfare.ie, resolved but never completed a TCP connection. One domain in our list, an-post.ie, no longer resolves at all; anpost.com is the live property and it served its file without complaint.

A 403 on robots.txt is permission, not protection

This is the part that surprises people. The Robots Exclusion Protocol was standardised as RFC 9309 in 2022, and section 2.3.1.3 covers exactly this case. If the server returns a status code in the 400 to 499 range, “the crawler MAY access any resources on the server”. Section 2.3.1.4 covers the opposite case: if the file is unreachable because of a server or network error, the crawler “MUST assume complete disallow”.

Read those two sentences together and the ranking is counter-intuitive. The single most protected site in our sample is www.welfare.ie, whose server never answered at all. The 17 sites returning 403 have, as far as a well-behaved crawler is concerned, published no rules whatsoever.

Four of those 403 responses carried Cloudflare’s cf-mitigated: challenge header, which tells you the mechanism: robots.txt is sitting behind the same managed challenge as the rest of the site. A human never notices, because a real browser solves the challenge in the background. No crawler can solve it. Others were Akamai, and a few were application-level: daft.ie, donedeal.ie and adverts.ie each returned a full HTML page, 73 KB to 94 KB of it, under a 403 status, for a file that should be about 400 bytes of plain text.

Five sites that answer a browser and refuse a crawler

We then repeated the request against all 76 working sites using the exact user agent string from OpenAI’s bot documentation, the variant with the robots.txt marker that OpenAI adds so site owners can tell those requests apart in their logs. Five sites flipped: aviva.ie, skillnetireland.ie, ucd.ie and gov.ie returned 403, and ul.ie returned 405.

The gov.ie case is worth dwelling on, because its robots.txt is one of the more carefully written in the sample. It disallows GPTBot, ClaudeBot, Google-Extended, Applebot, Applebot-Extended, Meta-ExternalAgent and Amazonbot, and it pointedly does not disallow OAI-SearchBot. Somebody thought about this. That thinking is delivered to a browser and withheld from the crawler it was written for.

One honest caveat: real OAI-SearchBot requests originate from published IP ranges, and most edge platforms treat verified bots differently from an unverified client sending the same string. Our requests came from an ordinary address, so a 403 may be bot verification working correctly rather than a blanket block. That is precisely the problem. Your crawling policy is now contingent on edge behaviour that nobody on the team has audited, for a file that has no reason to be gated at all. We have written before about why a user agent string is a claim rather than an identity, and verification is the right instinct. Applying it to robots.txt is not.

Most Irish sites have no AI policy at all

Of the 76 readable files, 57 do not name a single AI agent. No GPTBot, no ClaudeBot, no Google-Extended, nothing. That is not automatically wrong, because allowing everything is usually the right call for a site that wants to be cited. It is only a problem when it is an accident, which for 57 of 76 it almost certainly is.

The 19 that do say something are overwhelmingly media. Thirteen are news publishers. Only six are not: anpost.com, dcu.ie, gov.ie, isme.ie, myhome.ie and smythstoys.com. In a sample containing every major Irish bank, retailer and insurer, six non-media organisations have formed a view on AI crawling and written it down.

The blocklists are copied, and the copies are stale

Here is where it gets expensive. Anthropic’s own support documentation, updated 7 April 2026, lists three robots: ClaudeBot for training data, Claude-User for user-initiated fetches, and Claude-SearchBot for search indexing. In our sample, 10 files disallow anthropic-ai and 8 disallow claude-web. Neither token appears in that documentation. Exactly one file, independent.ie, names Claude-User and Claude-SearchBot.

The same pattern repeats elsewhere. Five files disallow perplexity-ai, a token that does not appear in Perplexity’s bot documentation, which lists PerplexityBot and Perplexity-User. These are not policies. They are copy-pasted lists that were current some time in 2024 and have not been revisited since, which is the same failure mode as an unmaintained dependency: it looks like it is doing something.

The live directives are pointing the wrong way

Read the vendor documentation carefully and the three jobs separate. OpenAI’s own words: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers.” Eight sites in our sample disallow it. Five of those are regional news brands whose AI blocks are identical token for token, the same 12 user agents in the same order, which tells you the decision was made once at group level and inherited.

Meanwhile five sites disallow ChatGPT-User, and three disallow Perplexity-User. Both vendors document these as user-initiated, and both say the rule may not bind. OpenAI: “Because these actions are initiated by a user, robots.txt rules may not apply.” Perplexity is blunter: “Since a user requested the fetch, this fetcher generally ignores robots.txt rules.” So the rules that are honoured remove you from AI answers, and several of the rules written to keep AI out are the ones least likely to be enforced.

Google-Extended is the most misread token of all. Google documents that it governs training for Gemini models and grounding, which is the step that supplies live Search content to the model at prompt time. Six sites disallow it, gov.ie among them. Public health and welfare guidance that is explicitly allowed into ChatGPT’s search index is excluded from Gemini’s grounded answers. gov.ie also disallows Applebot outright, and Apple is unambiguous that Applebot is what puts a site into Spotlight, Siri and Safari search. That is a search engine opt-out sitting in a list of AI opt-outs.

What to actually do

Four things, in order of how cheap they are.

Serve robots.txt unconditionally. Exempt the path from WAF rules, managed challenges, rate limiting and geo-blocking at the edge. It is a public text file containing instructions you want machines to read. Gating it achieves nothing except making your policy invisible to the only audience it has.

Monitor it like an endpoint. Add a synthetic check that requests /robots.txt with the documented user agent strings for GPTBot, OAI-SearchBot, ClaudeBot and Googlebot, and alerts on anything that is not a 200 with content type text/plain. Every site in our 403 group would have caught this on day one.

Make the three decisions separately. Training, search indexing and user-initiated retrieval are different products with different tokens and different commercial consequences. Decide each one on purpose, then write current token names. If you want out of training but into AI answers, that is GPTBot and Google-Extended disallowed, with OAI-SearchBot, Claude-SearchBot, PerplexityBot and Applebot allowed.

Verify by IP, not by string. OpenAI publishes four range files, and we pulled them while writing this: 18 prefixes for GPTBot, 39 for OAI-SearchBot, 230 for ChatGPT-User and 2 for OAI-AdsBot. The ChatGPT-User file was last regenerated on 25 September 2026, eleven days before this post. These are feeds, not a one-off copy into a firewall rule. If you are running ads inside ChatGPT, note that OAI-AdsBot visits only the landing pages you submit, from two prefixes, and an edge rule that blocks it will fail your ad review rather than your SEO.

The deadline on this is moving. Cloudflare shipped a Web Search API on 2 October 2026 that gives agents web results through Ceramic.ai, Exa and Linkup. In our 76 files, one names Exa. None name Linkup or Ceramic.ai. The list of fetchers your site needs an opinion about is growing faster than anyone’s robots.txt, which is an argument for a short, current, reachable file rather than a long inherited one.

REPTILEHAUS builds and audits this layer for clients: edge and WAF configuration, crawler and bot policy, AI visibility, and the monitoring that tells you when a CDN change quietly took your rules offline. If you are not certain what a crawler sees when it asks your site for instructions, that is a short engagement with a definite answer. Get in touch.

Method note: all figures are from first-hand HTTP requests made on 6 October 2026 against 100 Irish domains, with vendor behaviour quoted from OpenAI, Anthropic, Google, Apple and Perplexity’s current public documentation.

📷 Photo by Masaaki Komori on Unsplash