Anthropic documents three web agents, not one: ClaudeBot, Claude-User and Claude-SearchBot. Its support article on crawling asks site owners for a separate User-agent: block per agent, which means a rule you wrote for one of the three governs one of the three. This lesson gets you to the point of proving, against your own file, which group is in force for each name.
Then it takes the harder half. The IP prefix file Anthropic publishes so that owners can verify its agents lists network ranges and no agent names. Checked against it, a request is Anthropic’s or it is not; nothing says which of the three sent it. That answer lives only in the user-agent string, which is what a client says about itself. So “ClaudeBot hit us 4,000 times, Claude-User twice” is a tally of self-descriptions, and no published data upgrades it.
A step earlier, measured on this host: fetch /robots.txt from doctor-seo.net repeatedly on 30 September 2026 and some requests hand back the file while others hand back something else under the same HTTP 200. An audit that parses whatever arrived reports on the wrong document, which is why the diagnostic validates its inputs before judging a row.
Scope, twice over. Every vendor’s crawler in one table is Level 4 on GEO and AIO, not covered again here. And on the question readers arrive with, block them or not, this lesson recommends neither, sets out what each camp has, and says what would settle it.
What you’ll learn
- Why a
Disallowwritten for the ClaudeBot user agent leaves Claude-User and Claude-SearchBot exactly as they were. - What an address check establishes about a request and what it cannot, and the ceiling that puts on log analysis.
- How to audit your
robots.txtand access log offline, with fourteen single-cause exit codes and fixtures that reproduce every figure here. - How to find out whether your own
/robots.txtis even reachable as a file, with one standard-library script. - What each camp on blocking AI crawlers has, what none of them has, and the measurement that would end the argument.
Is your robots.txt even reachable as a file?
Ask for /robots.txt twice and you can be handed two different documents, both under HTTP 200. On doctor-seo.net, measured on 30 September 2026, some requests return the file — text/plain, 159 bytes, one distinct body across every hit, SHA-256 prefix 312e6b75a110 — and others roughly 12.3 KB of text/html, an edge verification interstitial wearing the same status code.
How often is the wrong question, and the numbers show why. Two runs of ten cache-busted requests, minutes apart on one host, returned the file 6 of 10 and then 7 of 10. The mechanism reproduced on every request and the proportion moved by a tenth from one run to the next, so what is established is that the substitution happens, not at what rate.
One variable moves it hard. Three requests sent with no User-Agent header, so the library’s own string went out, came back HTTP 403, three of three. What a crawler receives at that URL turns on what it calls itself, before any directive inside the file is read.
For an audit the damage is mechanical. Hand 12 KB of HTML to a robots.txt parser and it finds no user-agent field, concludes that nothing matches, and prints the most comforting output in its repertoire: no rules, nothing to fix. An empty file produces the same report. One standard-library script settles which document you are looking at:
#!/usr/bin/env python3
"""fetch-robots.py - fetch your own /robots.txt n times, cache-busted, and count
how often you got the file rather than something else carrying a 200.
Usage: python3 fetch-robots.py https://example.com 10 [named|stock]
named send User-Agent: robots-check/1.0 (the default)
stock send no User-Agent header, so the library's own string goes out
"""
import hashlib
import sys
import time
import urllib.error
import urllib.request
PAUSE = 1.2 # seconds between requests, so a run paces itself
base, count = sys.argv[1].rstrip("/"), int(sys.argv[2])
mode = sys.argv[3] if len(sys.argv) > 3 else "named"
headers = {} if mode == "stock" else {"User-Agent": "robots-check/1.0"}
tag = time.strftime("%Y%m%d%H%M%S")
print(f" mode={mode} n={count} host={base}")
ok, other, digests = 0, 0, set()
for i in range(count):
url = f"{base}/robots.txt?cb={tag}{i:02d}"
try:
response = urllib.request.urlopen(
urllib.request.Request(url, headers=headers), timeout=25)
body, status = response.read(), response.status
ctype = (response.headers.get("content-type") or "").split(";")[0]
except urllib.error.HTTPError as exc:
status, ctype, body = exc.code, "-", b""
if status == 200 and ctype == "text/plain":
ok += 1
digests.add(hashlib.sha256(body).hexdigest()[:12])
verdict = "the file"
else:
other += 1
verdict = "NOT the file"
print(f" {i:>2} {status} {ctype:<12} {len(body):>6} bytes {verdict}")
time.sleep(PAUSE)
print(f"\n {ok}/{count} returned text/plain; {other}/{count} did not.")
print(f" distinct text/plain bodies: {len(digests)} {sorted(digests)}")
Three runs, verbatim, as produced on 30 September 2026 — the two named-agent runs above and the stock-user-agent run:
$ python3 fetch-robots.py https://doctor-seo.net 10 named
mode=named n=10 host=https://doctor-seo.net
0 200 text/plain 159 bytes the file
1 200 text/plain 159 bytes the file
2 200 text/html 12315 bytes NOT the file
3 200 text/html 12330 bytes NOT the file
4 200 text/plain 159 bytes the file
5 200 text/html 12234 bytes NOT the file
6 200 text/plain 159 bytes the file
7 200 text/html 12306 bytes NOT the file
8 200 text/plain 159 bytes the file
9 200 text/plain 159 bytes the file
6/10 returned text/plain; 4/10 did not.
distinct text/plain bodies: 1 ['312e6b75a110']
$ python3 fetch-robots.py https://doctor-seo.net 10 named
mode=named n=10 host=https://doctor-seo.net
0 200 text/plain 159 bytes the file
1 200 text/html 12255 bytes NOT the file
2 200 text/plain 159 bytes the file
3 200 text/plain 159 bytes the file
4 200 text/html 12246 bytes NOT the file
5 200 text/plain 159 bytes the file
6 200 text/plain 159 bytes the file
7 200 text/plain 159 bytes the file
8 200 text/plain 159 bytes the file
9 200 text/html 12174 bytes NOT the file
7/10 returned text/plain; 3/10 did not.
distinct text/plain bodies: 1 ['312e6b75a110']
$ python3 fetch-robots.py https://doctor-seo.net 3 stock
mode=stock n=3 host=https://doctor-seo.net
0 403 - 0 bytes NOT the file
1 403 - 0 bytes NOT the file
2 403 - 0 bytes NOT the file
0/3 returned text/plain; 3/3 did not.
distinct text/plain bodies: 0 []
Two things it will not tell you. It cannot speak for what Anthropic’s agents receive, not being one of them. And one text/plain response is no proof the body arrived complete, since a truncated file with a plausible content type parses cleanly. Compare digests across runs; more than one is a question about your edge.
Three agents, and why one block covers one of them
Three tokens exist and a rule naming one of them names one of them. Anthropic’s support article, “Does Anthropic crawl data from the web, and how can site owners block the crawler?”, documents ClaudeBot, Claude-User and Claude-SearchBot and covers blocking them through robots.txt, a User-agent: block apiece. The file itself: plain text at a host’s root, declaring which paths each named program may request. How requesting differs from indexing and ranking is how Google works; the retrieval behind Claude’s answers is getting cited by Claude and the Brave dependency, and no citation figure appears on this page.
Being precise about what is citable here matters, because the mechanism gets asserted far more confidently than it gets sourced. The requirement of a block per agent is Anthropic’s, on the page linked above. That a block on one agent does not reach the others is Search Engine Land’s, 25 February 2026. What is not available to cite is a specification reference for the behaviour underneath both — that a group naming an agent replaces the wildcard rather than adding to it. Take substitution as the reading the per-agent requirement implies, and take the script below as what your own file demonstrates: it reports which group names each agent, a fact, and labels the consequence under that reading.
Play it forward and the commonest failure with these three falls out. Give ClaudeBot a group and ClaudeBot stops reading User-agent: *; put Disallow: / in that group and the site closes to ClaudeBot alone, while Claude-User and Claude-SearchBot carry on under whatever the wildcard said, or under nothing where no wildcard exists. Whoever wrote the file believed they had shut out Claude. They shut out a third of it, and the diagnostic prints that in one line.
Which of the three crawled you? The prefix file will not say
Anthropic publishes a prefix file at claude.com/crawling/bots.json so an owner can establish a request came from Anthropic. Read on 20 September 2026, it declared a creationTime of 18 August 2026 and carried 26 IPv4 prefixes over Google Cloud, Azure and AWS. Those are its contents: network ranges, a timestamp, and no agent name.
What follows is mechanical, and it concerns addresses rather than anybody’s file. An IP address identifies the network a request came from and nothing more; which piece of software sent it is asserted only by the request’s own headers. An address check is therefore the strong kind, because it takes the client’s word for nothing. But given a prefix list and no agent name, a log-based per-agent audit has nothing to check the user-agent string against, so no per-agent request count for these three is verified, whoever produced it.
One criticism here comes with its receipt. The support article owners are expected to act on renders no absolute last-updated date — a relative timestamp only, which a live check on 23 September 2026 established. For a page whose function is to say which agent tokens exist, that is a gap: nothing on it says whether three is still the number.
What each source settles, and does not:
| Source | Read on | Settles | Does not settle |
|---|---|---|---|
| Anthropic support article on crawling | Live check 23 Sep 2026 | Three agent names; a block for each | How current the list is |
claude.com/crawling/bots.json |
20 Sep 2026 | 26 IPv4 prefixes, a timestamp, no agent name | Which software sent a request from a listed address |
| The user-agent field in your log | Whenever you read it | What the client said about itself | What the client is |
Which group is really in force for each agent? Run this
claude-agents-audit.py reads three files off your disk and answers two questions about ClaudeBot, Claude-User and Claude-SearchBot: which group of robots.txt directives names each one, and how your log’s requests split between declared agents and published prefixes. It opens no socket, so the audit is pinned to a prefix file you kept.
Validation runs to completion before any row is judged, and that ordering is the design. A header dump, an HTML interstitial, an empty file, a log in the wrong format, a robots file cut mid-line and a prefix file cut mid-token each land on a code of their own, so none surfaces looking like a finding. Fourteen codes, one cause each:
| Code | Its single cause |
|---|---|
| 0 | All three agents are named by a group of their own |
| 1 | At least one agent is named by no group |
| 2 | The argument count is not three |
| 3 | A named file does not exist |
| 4 | A named file exists and will not open |
| 5 | The robots input is whitespace only |
| 6 | The robots input is an HTML document |
| 7 | The robots input is an HTTP header dump |
| 8 | The robots input has no user-agent field |
| 9 | A non-comment robots line carries no colon |
| 10 | The prefix file is whitespace only |
| 11 | The prefix file has content, no line parsed |
| 12 | The access log is whitespace only |
| 13 | The log has lines, none in combined format |
Code 1 claims less than it looks: a name is absent, not the file is wrong. Leaving all three on the wildcard is coherent, so the script hands you a fact, not a verdict.
#!/usr/bin/env python3
"""claude-agents-audit.py - which robots.txt group names each Anthropic agent,
and which log requests came from a published prefix. Three local files, no
network request of any kind.
Usage: python3 claude-agents-audit.py <robots.txt> <access.log> <prefixes.txt>
Fourteen exit codes. Each has exactly one cause, and no two share one. Every
input is validated before any row is judged, so no bad input can surface as a
finding about your configuration.
0 all three agents are named by a group of their own
1 at least one agent is named by no group in the file
2 the argument count is not three
3 a named file does not exist
4 a named file exists and cannot be read
5 the robots input is empty or whitespace only
6 the robots input is an HTML document
7 the robots input is an HTTP header dump
8 the robots input has no user-agent field
9 the robots input has a non-comment line carrying no colon
10 the prefix file is empty or whitespace only
11 the prefix file has content and no line parsed as a network prefix
12 the access log is empty or whitespace only
13 the access log has non-blank lines and none parsed as combined format
Cannot be detected, and stated so in the output: a user-agent string is chosen
by the client, so a declared agent is never verified; and a file truncated on a
clean line boundary parses as a shorter valid file, for the robots input and the
prefix file alike, which is why the counts are printed for you to check.
"""
import ipaddress
import os
import re
import sys
from collections import defaultdict
AGENTS = ("ClaudeBot", "Claude-User", "Claude-SearchBot")
LOGLINE = re.compile(
r'^(\S+) \S+ \S+ \[[^\]]*\] "[A-Z]+ [^"]*" \d{3} \S+ "[^"]*" "([^"]*)"$'
)
def die(code, message):
print(f"INPUT ERROR {code}: {message}", file=sys.stderr)
raise SystemExit(code)
def read(path):
if not os.path.exists(path):
die(3, f"{path} does not exist")
try:
with open(path, "rb") as handle:
return handle.read().decode("utf-8", "replace")
except OSError as exc:
die(4, f"{path} exists and cannot be read: {exc.strerror}")
def check_robots(text, path):
if not text.strip():
die(5, f"{path} is empty or whitespace only")
head = text.lstrip()[:400].lower()
if head.startswith("<!doctype") or head.startswith("<html") or "<head" in head:
die(6, f"{path} is an HTML document - a 200 with the wrong content type "
"is still not your robots.txt")
if re.match(r"^HTTP/[\d.]+ \d{3}", text.lstrip()):
die(7, f"{path} begins with a status line - this is a header dump")
if not re.search(r"(?im)^\s*user-agent\s*:", text):
die(8, f"{path} has no user-agent field anywhere in it")
for number, raw in enumerate(text.splitlines(), 1):
line = raw.strip()
if line and not line.startswith("#") and ":" not in line:
die(9, f"{path} line {number} carries no colon: {line!r} - "
"a malformed directive, or a file cut off mid-line")
def groups(text):
"""A run of consecutive user-agent lines shares one group of rules."""
out, current, collecting = [], None, False
for raw in text.splitlines():
line = raw.split("#", 1)[0].strip()
if not line or ":" not in line:
continue
field, _, value = line.partition(":")
field, value = field.strip().lower(), value.strip()
if field == "user-agent":
if not collecting:
current = {"names": [], "rules": []}
out.append(current)
collecting = True
current["names"].append(value.lower())
elif field in ("allow", "disallow") and current is not None:
collecting = False
current["rules"].append((field, value))
return out
def governing(parsed, agent):
named = [g for g in parsed if agent.lower() in g["names"]]
if named:
return named[0], "own group"
wildcard = [g for g in parsed if "*" in g["names"]]
if wildcard:
return wildcard[0], "wildcard * only"
return None, "no group at all"
def prefixes(text, path):
if not text.strip():
die(10, f"{path} is empty or whitespace only")
nets, rejected = [], 0
for raw in text.splitlines():
line = raw.strip().strip('",[] ')
if not line or line in ("{", "}"):
continue
if "/" not in line:
rejected += 1 # a prefix list holds prefixes, not addresses
continue
try:
# strict=True: host bits set means malformed, which is also what a
# line cut mid-token looks like (192.0.2.0/2 out of 192.0.2.0/25).
nets.append(ipaddress.ip_network(line, strict=True))
except ValueError:
rejected += 1
if not nets:
die(11, f"{path} has content and no line parsed as a network prefix "
f"({rejected} rejected)")
return nets, rejected
def listed(address, nets):
try:
ip = ipaddress.ip_address(address)
except ValueError:
return False
return any(ip in net for net in nets)
def declared(user_agent):
for agent in AGENTS:
if re.search(r"(?<![A-Za-z0-9-])" + re.escape(agent) + r"(?![A-Za-z0-9-])",
user_agent, re.I):
return agent
return None
def main(argv):
if len(argv) != 3:
die(2, "usage: claude-agents-audit.py <robots.txt> <access.log> <prefixes.txt>")
robots_path, log_path, prefix_path = argv
robots_text, log_text, prefix_text = (read(p) for p in argv)
check_robots(robots_text, robots_path)
nets, rejected = prefixes(prefix_text, prefix_path)
if not log_text.strip():
die(12, f"{log_path} is empty or whitespace only")
lines = [l for l in log_text.splitlines() if l.strip()]
rows = [(m.group(1), m.group(2)) for m in
(LOGLINE.match(l.strip()) for l in lines) if m]
if not rows:
die(13, f"{log_path} has {len(lines)} non-blank line(s) and none parsed "
"as combined log format")
print("== 1. Which group names each agent ==")
parsed = groups(robots_text)
unnamed = []
for agent in AGENTS:
group, source = governing(parsed, agent)
rules = "; ".join(f"{f.capitalize()}: {v}" for f, v in group["rules"]) if group else "-"
print(f" {agent:<17} {source:<17} {rules or '(group carries no rules)'}")
if source != "own group":
unnamed.append(agent)
print(f" groups parsed: {len(parsed)} | prefixes loaded: {len(nets)}"
f" | prefix lines rejected: {rejected}")
print()
print("== 2. Log requests: declared agent against published prefixes ==")
tally = defaultdict(lambda: [0, 0])
prefix_no_agent = 0
for address, user_agent in rows:
agent = declared(user_agent)
inside = listed(address, nets)
if agent:
tally[agent][0 if inside else 1] += 1
elif inside:
prefix_no_agent += 1
for agent in AGENTS:
inside, outside = tally[agent]
print(f" {agent:<17} {inside + outside:>4} declared |"
f" {inside:>4} from a listed prefix | {outside:>4} from elsewhere")
print(f" listed prefix, no agent declared: {prefix_no_agent}")
print(f" lines parsed: {len(rows)} of {len(lines)} non-blank")
print()
print("== 3. What section 2 does not establish ==")
print(" Not which agent sent any request. The published prefix list carries no")
print(" agent names, so the split above rests on a string the client chose.")
print(" Not that a request from elsewhere is not Anthropic's, and not that any")
print(" request obeyed a single line reported in section 1.")
print(" Not that these files arrived whole: truncation on a line boundary parses.")
if unnamed:
print()
print(f"UNNAMED: {', '.join(unnamed)} - named by no group in this file.")
print("A fact about the file. Whether it should be otherwise is your call.")
return 1
print()
print("NAMED: each of the three agents has a group of its own.")
return 0
if __name__ == "__main__":
raise SystemExit(main(sys.argv[1:]))
Point it at fixtures first. The generator writes every file the runs use. Two are not constructed: doctor-seo-robots.txt is the 159 bytes this host served on 30 September 2026, CRLF endings and lowercase fields intact, and robots-truncated.txt is its first twenty bytes. One prefix fixture is a /25, so prefix length has to be respected:
#!/bin/sh
# make-fixtures.sh - every file the runs below use. All constructed except
# doctor-seo-robots.txt, which is the 159 bytes doctor-seo.net served on
# 30 September 2026, and robots-truncated.txt, its first 20 bytes.
#
# The user-agent strings in access.log are SYNTHETIC PLACEHOLDERS wrapped around
# the documented agent tokens. No real Anthropic user-agent string is reproduced
# here or anywhere in this lesson: matching needs the token and nothing else.
set -e
cat > one-named.txt <<'EOF'
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
User-agent: ClaudeBot
Disallow: /
EOF
cat > three-named.txt <<'EOF'
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-User
Disallow: /wp-admin/
User-agent: Claude-SearchBot
Disallow: /wp-admin/
EOF
printf 'User-agent: * \r\ndisallow:*/feed/*\r\ndisallow:*/?s=search_term_string\r\ndisallow:*/amp\r\ndisallow:*/search/*\r\n\r\n\r\nsitemap: https://doctor-seo.net/sitemap_index.xml' > doctor-seo-robots.txt
dd if=doctor-seo-robots.txt of=robots-truncated.txt bs=1 count=20 2>/dev/null
cat > interstitial.txt <<'EOF'
<!DOCTYPE html>
<html lang="en">
<head><meta charset="utf8"><title>One moment, please...</title></head>
<body><div id="text">Please wait while your request is being verified...</div></body>
</html>
EOF
printf 'HTTP/2 200\r\ncontent-type: text/plain; charset=utf-8\r\ncache-control: max-age=600\r\n\r\n' > headers.txt
cat > no-agent.txt <<'EOF'
Disallow: /cart/
Sitemap: https://example.com/sitemap.xml
EOF
cat > prefixes.txt <<'EOF'
192.0.2.0/25
198.51.100.0/24
EOF
printf '192.0.2.0/2' > prefixes-truncated.txt
cat > access.log <<'EOF'
192.0.2.10 - - [29/Sep/2026:06:11:03 +0000] "GET /en/seo-course/ HTTP/1.1" 200 41210 "-" "synthetic fixture: ClaudeBot"
192.0.2.140 - - [29/Sep/2026:06:12:19 +0000] "GET /en/seo-course/ai-for-seo/ HTTP/1.1" 200 56962 "-" "synthetic fixture: ClaudeBot"
198.51.100.7 - - [29/Sep/2026:06:13:44 +0000] "GET /robots.txt HTTP/1.1" 200 159 "-" "synthetic fixture: Claude-User"
198.51.100.61 - - [29/Sep/2026:06:14:02 +0000] "GET /sitemap_index.xml HTTP/1.1" 200 9042 "-" "synthetic fixture: Claude-SearchBot"
203.0.113.9 - - [29/Sep/2026:06:15:30 +0000] "GET /en/seo-course/ HTTP/1.1" 200 41210 "-" "synthetic fixture: ClaudeBot"
198.51.100.90 - - [29/Sep/2026:06:16:55 +0000] "GET /en/blog/ HTTP/1.1" 200 18110 "-" "synthetic fixture: no agent declared"
2001:db8::52 - - [29/Sep/2026:06:17:40 +0000] "GET / HTTP/1.1" 200 12004 "-" "synthetic fixture: browser"
EOF
cat > access-wrongformat.log <<'EOF'
[29/Sep/2026:06:11:03] GET /en/seo-course/ 200 41210
[29/Sep/2026:06:12:19] GET /en/seo-course/ai-for-seo/ 200 56962
EOF
: > empty.txt
Run one is the configuration failure: one-named.txt shuts out ClaudeBot and says nothing of the others. The log fixture’s user-agent strings are synthetic placeholders around the documented tokens, not strings any agent sends:
$ python3 claude-agents-audit.py one-named.txt access.log prefixes.txt
== 1. Which group names each agent ==
ClaudeBot own group Disallow: /
Claude-User wildcard * only Disallow: /wp-admin/; Allow: /wp-admin/admin-ajax.php
Claude-SearchBot wildcard * only Disallow: /wp-admin/; Allow: /wp-admin/admin-ajax.php
groups parsed: 2 | prefixes loaded: 2 | prefix lines rejected: 0
== 2. Log requests: declared agent against published prefixes ==
ClaudeBot 3 declared | 1 from a listed prefix | 2 from elsewhere
Claude-User 1 declared | 1 from a listed prefix | 0 from elsewhere
Claude-SearchBot 1 declared | 1 from a listed prefix | 0 from elsewhere
listed prefix, no agent declared: 1
lines parsed: 7 of 7 non-blank
== 3. What section 2 does not establish ==
Not which agent sent any request. The published prefix list carries no
agent names, so the split above rests on a string the client chose.
Not that a request from elsewhere is not Anthropic's, and not that any
request obeyed a single line reported in section 1.
Not that these files arrived whole: truncation on a line boundary parses.
UNNAMED: Claude-User, Claude-SearchBot - named by no group in this file.
A fact about the file. Whether it should be otherwise is your call.
# exit code 1
Section 2 carries the trap the fixture was built around. ClaudeBot shows three declared requests, one from a listed prefix and two not. One calls itself ClaudeBot from 203.0.113.9, outside every listed range, so nothing attributes it to Anthropic. The other is 192.0.2.140, and 192.0.2.0/25 stops at .127. Anyone reading that as the 192.0.2 block had counted it verified.
Name all three, and only the exit code moves:
$ python3 claude-agents-audit.py three-named.txt access.log prefixes.txt
== 1. Which group names each agent ==
ClaudeBot own group Disallow: /
Claude-User own group Disallow: /wp-admin/
Claude-SearchBot own group Disallow: /wp-admin/
groups parsed: 4 | prefixes loaded: 2 | prefix lines rejected: 0
== 2. Log requests: declared agent against published prefixes ==
ClaudeBot 3 declared | 1 from a listed prefix | 2 from elsewhere
Claude-User 1 declared | 1 from a listed prefix | 0 from elsewhere
Claude-SearchBot 1 declared | 1 from a listed prefix | 0 from elsewhere
listed prefix, no agent declared: 1
lines parsed: 7 of 7 non-blank
== 3. What section 2 does not establish ==
Not which agent sent any request. The published prefix list carries no
agent names, so the split above rests on a string the client chose.
Not that a request from elsewhere is not Anthropic's, and not that any
request obeyed a single line reported in section 1.
Not that these files arrived whole: truncation on a line boundary parses.
NAMED: each of the three agents has a group of its own.
# exit code 0
Run three is this host’s own file, a worked example and not a model. One wildcard group, four closed path patterns, no AI agent named:
$ python3 claude-agents-audit.py doctor-seo-robots.txt access.log prefixes.txt
== 1. Which group names each agent ==
ClaudeBot wildcard * only Disallow: */feed/*; Disallow: */?s=search_term_string; Disallow: */amp; Disallow: */search/*
Claude-User wildcard * only Disallow: */feed/*; Disallow: */?s=search_term_string; Disallow: */amp; Disallow: */search/*
Claude-SearchBot wildcard * only Disallow: */feed/*; Disallow: */?s=search_term_string; Disallow: */amp; Disallow: */search/*
groups parsed: 1 | prefixes loaded: 2 | prefix lines rejected: 0
== 2. Log requests: declared agent against published prefixes ==
ClaudeBot 3 declared | 1 from a listed prefix | 2 from elsewhere
Claude-User 1 declared | 1 from a listed prefix | 0 from elsewhere
Claude-SearchBot 1 declared | 1 from a listed prefix | 0 from elsewhere
listed prefix, no agent declared: 1
lines parsed: 7 of 7 non-blank
== 3. What section 2 does not establish ==
Not which agent sent any request. The published prefix list carries no
agent names, so the split above rests on a string the client chose.
Not that a request from elsewhere is not Anthropic's, and not that any
request obeyed a single line reported in section 1.
Not that these files arrived whole: truncation on a line boundary parses.
UNNAMED: ClaudeBot, Claude-User, Claude-SearchBot - named by no group in this file.
A fact about the file. Whether it should be otherwise is your call.
# exit code 1
And the twelve input faults, each on its own code. Note code 9: the live file cut to twenty bytes exits as an input error, where before that check it exited 1 and read like a finding about an intact file.
$ python3 claude-agents-audit.py one-named.txt access.log
INPUT ERROR 2: usage: claude-agents-audit.py <robots.txt> <access.log> <prefixes.txt>
# exit code 2
$ python3 claude-agents-audit.py missing.txt access.log prefixes.txt
INPUT ERROR 3: missing.txt does not exist
# exit code 3
$ python3 claude-agents-audit.py . access.log prefixes.txt
INPUT ERROR 4: . exists and cannot be read: Is a directory
# exit code 4
$ python3 claude-agents-audit.py empty.txt access.log prefixes.txt
INPUT ERROR 5: empty.txt is empty or whitespace only
# exit code 5
$ python3 claude-agents-audit.py interstitial.txt access.log prefixes.txt
INPUT ERROR 6: interstitial.txt is an HTML document - a 200 with the wrong content type is still not your robots.txt
# exit code 6
$ python3 claude-agents-audit.py headers.txt access.log prefixes.txt
INPUT ERROR 7: headers.txt begins with a status line - this is a header dump
# exit code 7
$ python3 claude-agents-audit.py no-agent.txt access.log prefixes.txt
INPUT ERROR 8: no-agent.txt has no user-agent field anywhere in it
# exit code 8
$ python3 claude-agents-audit.py robots-truncated.txt access.log prefixes.txt
INPUT ERROR 9: robots-truncated.txt line 2 carries no colon: 'disa' - a malformed directive, or a file cut off mid-line
# exit code 9
$ python3 claude-agents-audit.py one-named.txt access.log empty.txt
INPUT ERROR 10: empty.txt is empty or whitespace only
# exit code 10
$ python3 claude-agents-audit.py one-named.txt access.log prefixes-truncated.txt
INPUT ERROR 11: prefixes-truncated.txt has content and no line parsed as a network prefix (1 rejected)
# exit code 11
$ python3 claude-agents-audit.py one-named.txt empty.txt prefixes.txt
INPUT ERROR 12: empty.txt is empty or whitespace only
# exit code 12
$ python3 claude-agents-audit.py one-named.txt access-wrongformat.log prefixes.txt
INPUT ERROR 13: access-wrongformat.log has 2 non-blank line(s) and none parsed as combined log format
# exit code 13
What it cannot detect, stated rather than implied. A spoofed user-agent string: a request calling itself ClaudeBot from a listed prefix counts as declared-and-listed, and section 3 of every run says so. Truncation on a clean line boundary, in either file: a prefix list cut after a whole prefix, or a robots file cut after a whole directive, parses as a shorter valid file and nothing flags it, which is why the prefix and line counts print. And whether any client obeyed section 1.
Monthly scheduling suits the agent workflow Claude Code for technical SEO sets up; logs and crawl behaviour are Level 2 on technical SEO.
Does a Disallow line actually stop anything?
By itself, no. A group for ClaudeBot, Claude-User or Claude-SearchBot is a request honoured at the client’s discretion; turning a download away takes an error status or an authentication challenge. The 403s in the first section are the contrast in miniature: an edge rule refused requests outright, which no line of text in a file does.
Three roles a Disallow gets mistaken for. It is not a noindex: refusing a request leaves untouched every copy already made. It is not an identity check, which is the argument above. And it is not durable as an address list, because the prefix file regenerates: copy one in June and by September it turns away retired addresses and admits new ones. Search Engine Land flagged IP blocking as unreliable in the same 25 February 2026 article, no sample size.
Under the blocking question sits a prior one, live rather than neglected. OpenAI’s crawler documentation revision of 9 December 2025, reported by PPC Land, removed the statement that its user-initiated fetcher complies with robots.txt and put this in its place: “because these actions are initiated by a user, robots.txt rules may not apply.” Whether a fetch a person asked for is governed by robots.txt at all is therefore already open at one vendor, and this lesson gives no verdict on it. Nobody has published a log-level measurement of which user agent such a task presents. What hangs on it is whether a Disallow aimed at a user-initiated fetcher is a lever or a note of intent.
Writing the three blocks, whichever way you decide
Two shapes for ClaudeBot, Claude-User and Claude-SearchBot, and this lesson endorses neither. Shutting the site to all three means a group per exact token, each carrying the same closure:
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-User
Disallow: /
User-agent: Claude-SearchBot
Disallow: /
Now the version that catches people out. Putting Allow: / under the three names, to place a welcome on the record, is not the harmless annotation it looks like. A group that names an agent takes that agent off the wildcard, so whatever the wildcard was closing for it stops being closed. A cart path and an account path shut under * stand open to all three the moment three permissive named groups exist. Carried forward properly:
User-agent: *
Disallow: /cart/
Disallow: /my-account/
User-agent: ClaudeBot
Disallow: /cart/
Disallow: /my-account/
User-agent: Claude-User
Disallow: /cart/
Disallow: /my-account/
User-agent: Claude-SearchBot
Disallow: /cart/
Disallow: /my-account/
Either way, re-run the audit. Its first section names the group that ended up in force for each agent, the only dependable way to tell what the file says from what you meant it to say.
Is blocking the right call, and who has measured it?
Nobody has measured it, and that is the defensible answer. Whether to shut out ClaudeBot is a decision about what your writing is worth and what you will accept for it; three blocks is the easy half. Here is each position’s best evidence and the hole all of them share.
The figures everybody quotes come from Cloudflare Radar’s snapshot of 31 August 2026, across the top roughly 10,000 domains, of which some 4,000 robots.txt files parsed. Per crawler: ClaudeBot 619 Disallow against 259 Allow, 2.39:1; GPTBot 696 against 299, 2.33:1; OAI-SearchBot 250 against 266, 0.94:1, a net allow. Sorting those crawlers into training and answering agents is Cloudflare Radar’s classification.
Read one way that spread is settled practice: publishers shut out training agents two to three times as often as they admit them and net-admit the ones that return a link, so the market has drawn the distinction and a file that names nobody has made no decision. Read the other way, those columns count lines in text files, not a visit or a citation or a pound of revenue on either side of a block. A third position argues training access belongs under licence and should be charged for, and has momentum rather than measurement: Cloudflare changed its own crawler defaults on 15 September 2026, while no adoption or revenue figures for metered crawling have been published.
Two limits, the second of which this page adds. The frame: ten thousand domains are not the web, and a count of directives is not an outcome, so those ratios describe a slice of publisher behaviour and predict nothing. And the verification gap belongs to neither camp: it argues neither for blocking nor against it, but limits what both may claim, because a per-agent policy whose effects nobody can attribute to an agent cannot be shown to have worked or failed.
So this lesson recommends neither blocking nor allowing. What would end the argument is a study nobody has run: two matched sets of comparable sites, one blocking and one not, reporting citations and referred traffic on both sides across a stated window, sample size disclosed. Until it exists, “blocking removes you from AI answers” and “blocking costs nothing” are one guess pointed two ways. What Level 8 on using the models as tools hands you instead is a file recording the decision you did make, and a script proving it records it.
Common mistakes
- Reading a 200 as “I have the file”. HTML served at
/robots.txtparses as a file with no rules, which looks like an all-clear; on this host on 30 September 2026 it arrived on a minority of requests in one run and a different minority in the next. Fix: check content type and byte count before parsing, and read any input-error exit as a fetch problem, not a configuration finding. - Adding
Allow: /to be friendly. A named group replaces the wildcard, so a permissive block opens whatever*was closing for that agent. Fix: copy every wildcard rule you mean to keep into each named group, then re-run the audit. - Publishing a per-agent split as though it were measured. The prefix file carries no agent names, so the division is self-reported. Fix: call it declared traffic, and state which half is verified.
- Letting a prefix length slide.
192.0.2.0/25is not192.0.2.0/24, and the file regenerates besides. Fix: download it again on audit day and let a library decide containment.
The short version
- Anthropic documents three agents — ClaudeBot, Claude-User and Claude-SearchBot — and its support article asks for a separate
User-agent:block for each. - A group naming an agent takes that agent off
User-agent: *, soDisallow: /under ClaudeBot leaves the other two where they were andAllow: /under a named agent reopens what the wildcard was closing. No specification reference for that substitution is cited here: the per-agent requirement is Anthropic’s, the consequence for the other two is Search Engine Land’s, 25 February 2026. - Anthropic’s prefix file at
claude.com/crawling/bots.json, read 20 September 2026, held 26 IPv4 prefixes, acreationTimeof 18 August 2026 and no agent names. - So a listed address settles that a request is Anthropic’s and never which agent sent it, which makes every per-agent count self-reported, whoever published it.
- Requests to doctor-seo.net’s own
/robots.txton 30 September 2026 returned the file on some attempts and a 12.3 KB HTML interstitial under the same HTTP 200 on others, 6 of 10 in one run and 7 of 10 in the next, so the failure is established and its frequency is not. - Whether to block is unresolved: Cloudflare Radar’s 31 August 2026 snapshot of the top roughly 10,000 domains counts directives, not outcomes, and no study compares matched sites with and without a block.
Frequently asked questions
What does the ClaudeBot user agent string look like in my logs?
Matching turns on the token, ClaudeBot, and that is what a group has to name. The three tokens are documented; no complete user-agent string from Anthropic appears on this page, because none could be confirmed; the log fixture carries synthetic placeholders around the tokens, not captured strings. Whatever trails the token plays no part in matching, so a rule or log filter built around one full string you happened to capture is built on sand.
Can I verify Anthropic’s crawlers with a reverse DNS lookup, the way I verify Googlebot?
Not on anything available to cite. What Anthropic publishes for verification is the prefix file at claude.com/crawling/bots.json — 26 IPv4 prefixes and no agent names when read on 20 September 2026 — and no forward-and-reverse DNS procedure for these agents could be confirmed. The check you have is address containment, which establishes the operator rather than the agent.
ClaudeBot is still requesting pages after I added a Disallow. Why?
Three candidates, and they separate cleanly. Your group may not spell the token exactly, leaving ClaudeBot on the wildcard, which the diagnostic reports in a line. Or robots.txt is honoured at the client’s discretion, and a client that declines contradicts nothing a text file can compel. Or the requests are not ClaudeBot: put the addresses against the published prefixes. Cheapest first.
Sources
- Linkable. Anthropic, “Does Anthropic crawl data from the web, and how can site owners block the crawler?” — documents the three agents and blocking through
robots.txt. It renders a relative timestamp and no absolute date, established by a live check on 23 September 2026, so no publication date is quoted. - Vendor files and documentation, cited without a link because neither could be confirmed. Anthropic,
claude.com/crawling/bots.json, read 20 September 2026:creationTime18 August 2026, 26 IPv4 prefixes, no agent names. OpenAI, crawler documentation revision of 9 December 2025, reported by PPC Land: the statement that its user-initiated fetcher complies withrobots.txtgave way to “because these actions are initiated by a user, robots.txt rules may not apply”. - Trade press, cited without a link for the same reason. Search Engine Land on Anthropic’s crawlers, 25 February 2026: a block on one agent does not reach the others, and IP blocking is unreliable. Documentation reporting, no sample size.
- Third-party measurement, cited without a link. Cloudflare Radar, robots.txt snapshot of 31 August 2026, top roughly 10,000 domains, some 4,000 files parsed: ClaudeBot 619 Disallow / 259 Allow (2.39:1), GPTBot 696 / 299 (2.33:1), OAI-SearchBot 250 / 266 (0.94:1). The training-versus-answering classification is Cloudflare Radar’s. Cloudflare’s own defaults changed on 15 September 2026.
- Demand evidence. Google Autocomplete probe, 20 September 2026: the seed
claudebotreturnedclaudebot crawlerandclaudebot user agent. Demand present, no volume. - First-party measurement. doctor-seo.net, 30 September 2026: two runs of ten cache-busted GETs of
/robots.txtunder the user-agentrobots-check/1.0returned HTTP 200text/plainat 159 bytes on 6 and on 7 of 10, one distinct body throughout, SHA-256 prefix312e6b75a110, the rest HTTP 200text/htmlat roughly 12.3 KB; three requests sending noUser-Agentheader returned HTTP 403. Twenty-three requests, one host, one day. - Own material.
fetch-robots.py,claude-agents-audit.pyandmake-fixtures.sh, executed 30 September 2026. Standard library, fourteen exit codes.
Continue the course
- Related in this level: Getting Cited by Claude: The Brave Dependency, Being Findable by Brave’s Web Discovery Project, Claude Code for Technical SEO.
- Level: AI for SEO: Using the Models as Tools
- Course index: SEO course