Skip to content
Doctor SEO

AI for SEO: Using the Models as Tools

Getting Cited by Claude: The Brave Dependency

Claude reaches the web through Brave Search, and the two overlap figures the industry quotes rest on 15 queries and no published sample size. Run the script and fixtures here to produce your own number, with its denominator and its n attached.

Updated 1 October 2026 34 min read

By the end of this lesson you will be able to produce your own overlap figure between what Claude cites and what a search engine ranks, denominator and sample size beside it, and to say why the two figures the industry quotes for Claude are not comparable with each other.

The number repeated most often: Profound reported on 21 March 2025 an 86.7% overlap between the pages Claude cited and Brave Search’s top non-sponsored results, at p<0.0001, on 15 queries. Profound publishes that percentage and sample size; the rest of this paragraph is doctor-seo.net dividing. Multiply out: 13 hits, each query worth 6.7 points, 95% Wilson interval 62.1% to 96.3%.

The number quoted beside it counts something else. MERJ reported in 2026, no month published, that 79.2% of the URLs Claude cited sat in Brave’s top 10 against 34% for Google’s top 10, publishing no sample size at all. Profound’s denominator is queries, MERJ’s is cited URLs.

What follows is an instrument, which is why it lives in Level 8, AI for SEO: the subject is measuring which index one assistant draws from. Turning a figure into advice about page content, for AI search engines generally, is Level 4, GEO and AIO, and that division gets a section below.

What you’ll learn

  • Read any overlap percentage by naming its denominator first.
  • Run a script over two files you captured yourself for an overlap figure with its n, its interval and a per-query breakdown.
  • Watch identical data return 25.0%, 62.5% or 75.0% on nothing but when two URLs count as one.
  • State what Anthropic’s subprocessor listing does and does not establish.
  • Say what this evidence supports and which questions remain unsettled.

Two overlap percentages for Claude, two denominators, no comparison

An overlap percentage means nothing until you know what was counted, and the two figures published about Claude count different objects. Profound’s 86.7% counts queries: 15 asked, 13 matching Brave’s top non-sponsored results. MERJ’s 79.2% counts cited URLs sitting in Brave’s top 10. A hit rate per question asked against a hit rate per link returned, so an answer citing six URLs adds one unit to the first and six to the second.

MERJ’s pair carries a second trap, and it is why the pair beats either half. The 79.2% belongs to Brave’s top 10, the 34% to Google’s top 10, and either number without its engine cannot be checked. The contrast is the informative part, holding the method fixed and swapping the index: the only controlled comparison here. The missing denominator is what it does not survive, no count of URLs meaning no interval around either figure.

Profound also reports 20% for ChatGPT, and there are two ways to mishandle it. Taken at face value that 20% is the overlap between what Claude cites and what ChatGPT cites; no comparison set is recorded, so it is not ChatGPT measured against an index of its own. Standing the two numbers side by side as one method run on two engines invents a methodology Profound never describes. Dropping the 20% is the opposite error, a summary omitting a figure the study reports not being that study. A third, widely circulated collection of Claude-cited URLs gets no percentage here at all, its data and collection dates being untraceable: a figure to one decimal place is not a stronger claim than one with a sample size.

Source and date What the percentage counts Figure Sample size published What it cannot support
Profound, 21 March 2025 Queries where Claude’s citations matched Brave’s top non-sponsored results 86.7% at p<0.0001; also 20% for ChatGPT Yes: 15 queries Any claim about a population, or any pairing with the 20%, which records no comparison set
MERJ, 2026, no month published URLs cited by Claude present in one engine’s top 10 79.2% for Brave’s top 10; 34% for Google’s top 10 No Any interval, having no denominator. The contrast survives, the precision does not
A widely circulated dataset of Claude-cited URLs Cited URLs by top-level domain and path Not reproduced here A URL count, no collection dates Anything: neither the data nor its collection window was locatable

Those two rows are the whole published evidence base for the claim that Claude’s citations track Brave Search: one with 15 queries behind it, one with no stated sample. Which argues for producing your own number.

Measure your own overlap instead of quoting somebody else’s

overlap.py takes two tab-separated files you capture yourself and prints an overlap figure with its denominator, its n, its 95% interval and a per-query breakdown. cited.tsv holds one line per URL an assistant cited, with the query behind it; ranked.tsv one line per result captured from a single engine, with its rank. It makes no network requests.

It prints two percentages, being the two that get confused. Cited-side divides matches by the cited URLs judged; ranked-side divides by the ranked URLs judged. They move independently, and reporting one without saying which is what makes published figures unreadable.

Two switches change the answer, which is the point. --top N sets how deep a match may be found, default 10. --match decides when two URLs count as one: exact compares strings as written; normalised lower-cases the host, drops www., the query and the fragment, ignores a trailing slash; host compares hosts alone. No published study names its rule.

Nothing is judged until both files have been read through and checked. A line that does not parse is dropped and the count printed, so a half-broken export cannot pass as a whole one. The three ways this measurement quietly yields a beautiful zero exit on three separate numbers: a header row with no data under it, a file in the wrong format, a capture taken for different queries. Eleven codes, no two describing the same fault:

Exit The one thing it means
0 A figure was computed and printed.
2 The wrong number of file arguments arrived.
3 An option the script does not accept; it takes --top and --match only.
4 A file named on the command line could not be opened.
5 --top got something other than a positive integer.
6 --match got a rule name the script does not have.
7 The citation file carries no data line: empty, or only blanks, comments, a header.
8 It carries data lines, none of them query, tab, URL.
9 The results file carries no data line.
10 It carries data lines, none of them query, tab, rank, tab, URL.
11 The files share no query at --top N: nothing to divide by.

What the script cannot detect belongs here rather than in a footnote. It cannot tell you whether your citation capture is complete, since the same question asked twice may cite a different set. It cannot tell you whether your captured results are what an assistant’s retrieval saw when it saw them: locale, session and timing all move a result page. It fetches nothing, so it cannot know a URL still resolves, and it infers nothing about which index was queried. A high figure is co-occurrence, never a mechanism.

#!/usr/bin/env python3
# overlap.py - work out a citation-overlap percentage from captures you took
# yourself, and print the denominator the percentage belongs to.
#
# Two tab-separated files. Blank lines, '#' comments and one header row are
# ignored in each.
#   cited.tsv    query <TAB> url                  URLs an assistant cited
#   ranked.tsv   query <TAB> rank <TAB> url       results you captured from one engine
#
#   python3 overlap.py cited.tsv ranked.tsv --top 10 --match normalised
#
# TWO percentages come out, because they answer different questions and get
# quoted as though they were one:
#   CITED-SIDE   hits / cited URLs judged    "how much of what it cited was ranked"
#   RANKED-SIDE  hits / ranked URLs judged   "how much of what was ranked got cited"
#
# Exit codes. One cause per code, and nothing wrong with the inputs shares a
# code with a number this script worked out.
#   0   a figure was computed and printed
#   2   the wrong number of file arguments arrived
#   3   an option this script does not accept (it accepts --top and --match, nothing else)
#   4   a file named on the command line could not be opened
#   5   --top was handed something other than a positive integer
#   6   --match was handed a rule name this script does not have
#   7   the citation file carries no data line: empty, or only blanks, comments and a header
#   8   the citation file carries data lines, and none of them is query<TAB>url
#   9   the results file carries no data line
#   10  the results file carries data lines, and none is query<TAB>rank<TAB>url
#   11  the two files name no query in common at --top N, so the denominator would be zero

import math
import os
import re
import sys
from urllib.parse import urlsplit

MATCHES = ("exact", "normalised", "host")

USAGE = ("usage: overlap.py <cited.tsv> <ranked.tsv> [--top N] [--match RULE]\n"
         "  cited.tsv   query <TAB> url\n"
         "  ranked.tsv  query <TAB> rank <TAB> url\n"
         "  --top N     keep ranks 1..N only; default 10\n"
         "  --match     exact | normalised | host; default normalised\n"
         "  --top and --match are the only options; -h and --help are not accepted")


def die(code, message):
    sys.stderr.write(message.rstrip() + "\n")
    sys.exit(code)


def slurp(path):
    if not os.path.isfile(path):
        die(4, "input error: cannot open %s - no such readable file" % path)
    try:
        with open(path, "r", encoding="utf-8", errors="replace") as handle:
            return handle.read().splitlines()
    except OSError as err:
        die(4, "input error: cannot open %s (%s)" % (path, err.strerror))


def candidates(lines, header):
    """Lines that are neither blank, nor a comment, nor the header row."""
    kept = []
    for line in lines:
        if not line.strip() or line.lstrip().startswith("#"):
            continue
        fields = [f.strip().lower() for f in line.split("\t")]
        if tuple(fields[:len(header)]) == header:
            continue
        kept.append(line)
    return kept


def key_for(url, rule):
    parts = urlsplit(url.strip())
    host = (parts.netloc or "").lower().split("@")[-1]
    host = re.sub(r":(80|443)$", "", host)
    host = re.sub(r"^www\.", "", host)
    if rule == "host":
        return host
    if rule == "exact":
        return url.strip()
    path = re.sub(r"/{2,}", "/", parts.path) or "/"
    if len(path) > 1:
        path = path.rstrip("/")
    return host + path


def read_cited(path):
    lines = candidates(slurp(path), ("query", "url"))
    if not lines:
        die(7, "no data in %s: every line is blank, a comment or the header row" % path)
    rows, bad = [], 0
    for line in lines:
        fields = line.split("\t")
        if len(fields) < 2 or not fields[0].strip() or not fields[1].strip():
            bad += 1
            continue
        rows.append((fields[0].strip(), fields[1].strip()))
    if not rows:
        die(8, "wrong format in %s: %d data line(s), none of them query<TAB>url - check "
               "the file is tab-separated" % (path, bad))
    return rows, bad


def read_ranked(path):
    lines = candidates(slurp(path), ("query", "rank", "url"))
    if not lines:
        die(9, "no data in %s: every line is blank, a comment or the header row" % path)
    rows, bad = [], 0
    for line in lines:
        fields = line.split("\t")
        if len(fields) < 3 or not fields[0].strip() or not fields[2].strip():
            bad += 1
            continue
        try:
            rank = int(fields[1].strip())
        except ValueError:
            bad += 1
            continue
        if rank < 1:
            bad += 1
            continue
        rows.append((fields[0].strip(), rank, fields[2].strip()))
    if not rows:
        die(10, "wrong format in %s: %d data line(s), none of them query<TAB>rank<TAB>url "
                "- check the file is tab-separated and the rank is a number" % (path, bad))
    return rows, bad


def wilson(hits, n, z=1.96):
    """95% interval for a proportion. Returns (low, high) as percentages."""
    p = hits / n
    d = 1 + z * z / n
    centre = (p + z * z / (2 * n)) / d
    margin = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
    return 100 * (centre - margin), 100 * (centre + margin)


def report(label, question, hits, n):
    low, high = wilson(hits, n)
    print("%s" % label)
    print("  question   %s" % question)
    print("  figure     %.1f%%  = %d of %d" % (100 * hits / n, hits, n))
    print("  95%% CI     %.1f%% to %.1f%%   (width %.1f points)" % (low, high, high - low))
    print("  one URL    moves the figure %.1f points" % (100.0 / n))
    print()


def main():
    argv = sys.argv[1:]
    top, rule = 10, "normalised"

    if "--top" in argv:
        at = argv.index("--top")
        if at + 1 >= len(argv):
            die(5, "--top needs a positive integer after it\n" + USAGE)
        try:
            top = int(argv[at + 1])
        except ValueError:
            die(5, "--top got %r, which is not an integer\n%s" % (argv[at + 1], USAGE))
        if top < 1:
            die(5, "--top got %r, which is not positive\n%s" % (argv[at + 1], USAGE))
        argv = argv[:at] + argv[at + 2:]

    if "--match" in argv:
        at = argv.index("--match")
        if at + 1 >= len(argv):
            die(6, "--match needs one of %s after it\n%s" % (", ".join(MATCHES), USAGE))
        rule = argv[at + 1]
        if rule not in MATCHES:
            die(6, "--match got %r; the rules are %s\n%s"
                % (rule, ", ".join(MATCHES), USAGE))
        argv = argv[:at] + argv[at + 2:]

    unknown = [a for a in argv if a.startswith("-")]
    if unknown:
        die(3, "%r is not an option this script accepts\n%s" % (unknown[0], USAGE))
    if len(argv) != 2:
        die(2, "expected 2 file arguments, got %d\n%s" % (len(argv), USAGE))

    # Both files are read and checked through before any row is judged.
    cited, cited_bad = read_cited(argv[0])
    ranked, ranked_bad = read_ranked(argv[1])

    kept = [r for r in ranked if r[1] <= top]
    by_query = {}
    for query, _, url in kept:
        by_query.setdefault(query, set()).add(key_for(url, rule))

    judged_queries = sorted({q for q, _ in cited if q in by_query})
    if not judged_queries:
        die(11, "no query in %s appears in %s at --top %d, so there is nothing to divide "
                "by" % (argv[0], argv[1], top))

    dropped_queries = sorted({q for q, _ in cited if q not in by_query})

    cited_n = cited_hits = 0
    per_query = []
    for query in judged_queries:
        urls = [u for q, u in cited if q == query]
        hits = sum(1 for u in urls if key_for(u, rule) in by_query[query])
        cited_n += len(urls)
        cited_hits += hits
        per_query.append((query, hits, len(urls), len(by_query[query])))

    cited_keys = {}
    for query, url in cited:
        if query in by_query:
            cited_keys.setdefault(query, set()).add(key_for(url, rule))
    ranked_n = ranked_hits = 0
    for query in judged_queries:
        ranked_n += len(by_query[query])
        ranked_hits += len(by_query[query] & cited_keys.get(query, set()))

    print()
    print("matching rule      %s" % rule)
    print("rank cut-off       top %d" % top)
    print("queries in cited   %d" % len({q for q, _ in cited}))
    print("queries judged     %d  (a cited query with no captured results is excluded)"
          % len(judged_queries))
    if dropped_queries:
        print("queries excluded   %d: %s" % (len(dropped_queries), ", ".join(dropped_queries)))
    if cited_bad or ranked_bad:
        print("lines dropped      %d in cited, %d in ranked: unparseable, and in no "
              "denominator below" % (cited_bad, ranked_bad))
    print()

    report("CITED-SIDE", "how much of what was cited was also ranked", cited_hits, cited_n)
    report("RANKED-SIDE", "how much of what was ranked was also cited", ranked_hits, ranked_n)

    print("per query")
    for query, hits, urls, results in per_query:
        print("  %-34s %d of %-3d cited URLs ranked   (%d results captured)"
              % (query, hits, urls, results))
    print()
    print("Both percentages describe these %d queries and nothing else. Quote the "
          "figure\nwith its denominator and its n, or it cannot be read."
          % len(judged_queries))
    sys.exit(0)


if __name__ == "__main__":
    main()

Four runs on shipped fixtures: 25.0%, 62.5% and 75.0% from one dataset

The runs below use constructed fixtures, not real captures, so anyone copying the script and the generator out of this page reproduces the output exactly. No assistant and no engine was queried to make them. Run the generator in an empty directory.

#!/bin/sh
# make-overlap-fixtures.sh - writes every file the runs in this lesson use:
# two good captures, and five broken ones for the failure paths.
# All CONSTRUCTED. No assistant and no engine was queried to make them.
# Replace the two good ones with your own captures and the arithmetic is yours.

cat > cited.tsv <<'TSV'
query	url
crawl budget explained	https://example.com/technical/crawl-budget/
crawl budget explained	https://example.org/guides/crawl-budget
crawl budget explained	https://docs.example.net/crawling/budget/
redirect chain fix	https://example.org/guides/redirects
redirect chain fix	https://example.com/technical/redirect-chains/
redirect chain fix	https://blog.example.io/2026/redirect-hops/
canonical tag rules	https://example.com/technical/canonical/
canonical tag rules	https://help.example.com/canonical-tags
log file analysis seo	https://example.org/guides/log-files
TSV

cat > ranked.tsv <<'TSV'
query	rank	url
crawl budget explained	1	https://example.com/technical/crawl-budget
crawl budget explained	2	https://www.example.org/guides/crawl-budget/
crawl budget explained	3	https://other.example/crawl/
crawl budget explained	4	https://another.example/budget/
crawl budget explained	5	https://docs.example.net/crawling/
redirect chain fix	1	https://example.org/guides/redirects/
redirect chain fix	2	https://example.com/technical/redirect-chains/
redirect chain fix	3	https://third.example/hops/
canonical tag rules	1	https://example.com/technical/canonical/
canonical tag rules	2	https://fourth.example/canonical/
canonical tag rules	11	https://help.example.com/canonical-tags
TSV

# Exit 7: the header row survived the copy and the data did not.
printf 'query\turl\n' > cited-header-only.tsv

# Exit 8: comma-separated, so no line is query<TAB>url.
cat > cited-commas.csv <<'CSV'
query,url
crawl budget explained,https://example.com/technical/crawl-budget/
CSV

# Exit 9: nothing at all in the file.
: > ranked-empty.tsv

# Exit 10: an HTTP header dump saved over the results capture.
printf 'HTTP/2 200 \r\ncontent-type: text/html; charset=utf-8\r\n\r\n' > ranked-headers.tsv

# Exit 11: a real capture, taken for queries the citation file never mentions.
cat > ranked-elsewhere.tsv <<'TSV'
hreflang implementation	1	https://example.com/international/hreflang/
hreflang implementation	2	https://example.org/guides/hreflang
TSV

# Exit 0 with lines dropped: three malformed rows among the good ones.
cat > cited-malformed.tsv <<'TSV'
query	url
crawl budget explained	https://example.com/technical/crawl-budget/
crawl budget explained,https://example.org/guides/crawl-budget
redirect chain fix
redirect chain fix	https://example.com/technical/redirect-chains/
	https://example.com/orphan/
canonical tag rules	https://example.com/technical/canonical/
TSV

The fixtures put the same pages in both files under different URL strings: a missing trailing slash, a missing www., and a shared host reached by a different path.

The four runs, verbatim, executed on 30 September 2026:

$ python3 overlap.py cited.tsv ranked.tsv

matching rule      normalised
rank cut-off       top 10
queries in cited   4
queries judged     3  (a cited query with no captured results is excluded)
queries excluded   1: log file analysis seo

CITED-SIDE
  question   how much of what was cited was also ranked
  figure     62.5%  = 5 of 8
  95% CI     30.6% to 86.3%   (width 55.7 points)
  one URL    moves the figure 12.5 points

RANKED-SIDE
  question   how much of what was ranked was also cited
  figure     50.0%  = 5 of 10
  95% CI     23.7% to 76.3%   (width 52.7 points)
  one URL    moves the figure 10.0 points

per query
  canonical tag rules                1 of 2   cited URLs ranked   (2 results captured)
  crawl budget explained             2 of 3   cited URLs ranked   (5 results captured)
  redirect chain fix                 2 of 3   cited URLs ranked   (3 results captured)

Both percentages describe these 3 queries and nothing else. Quote the figure
with its denominator and its n, or it cannot be read.
# exit code 0

$ python3 overlap.py cited.tsv ranked.tsv --match exact

matching rule      exact
rank cut-off       top 10
queries in cited   4
queries judged     3  (a cited query with no captured results is excluded)
queries excluded   1: log file analysis seo

CITED-SIDE
  question   how much of what was cited was also ranked
  figure     25.0%  = 2 of 8
  95% CI     7.1% to 59.1%   (width 51.9 points)
  one URL    moves the figure 12.5 points

RANKED-SIDE
  question   how much of what was ranked was also cited
  figure     20.0%  = 2 of 10
  95% CI     5.7% to 51.0%   (width 45.3 points)
  one URL    moves the figure 10.0 points

per query
  canonical tag rules                1 of 2   cited URLs ranked   (2 results captured)
  crawl budget explained             0 of 3   cited URLs ranked   (5 results captured)
  redirect chain fix                 1 of 3   cited URLs ranked   (3 results captured)

Both percentages describe these 3 queries and nothing else. Quote the figure
with its denominator and its n, or it cannot be read.
# exit code 0

$ python3 overlap.py cited.tsv ranked.tsv --match host

matching rule      host
rank cut-off       top 10
queries in cited   4
queries judged     3  (a cited query with no captured results is excluded)
queries excluded   1: log file analysis seo

CITED-SIDE
  question   how much of what was cited was also ranked
  figure     75.0%  = 6 of 8
  95% CI     40.9% to 92.9%   (width 51.9 points)
  one URL    moves the figure 12.5 points

RANKED-SIDE
  question   how much of what was ranked was also cited
  figure     60.0%  = 6 of 10
  95% CI     31.3% to 83.2%   (width 51.9 points)
  one URL    moves the figure 10.0 points

per query
  canonical tag rules                1 of 2   cited URLs ranked   (2 results captured)
  crawl budget explained             3 of 3   cited URLs ranked   (5 results captured)
  redirect chain fix                 2 of 3   cited URLs ranked   (3 results captured)

Both percentages describe these 3 queries and nothing else. Quote the figure
with its denominator and its n, or it cannot be read.
# exit code 0

$ python3 overlap.py cited.tsv ranked.tsv --top 3

matching rule      normalised
rank cut-off       top 3
queries in cited   4
queries judged     3  (a cited query with no captured results is excluded)
queries excluded   1: log file analysis seo

CITED-SIDE
  question   how much of what was cited was also ranked
  figure     62.5%  = 5 of 8
  95% CI     30.6% to 86.3%   (width 55.7 points)
  one URL    moves the figure 12.5 points

RANKED-SIDE
  question   how much of what was ranked was also cited
  figure     62.5%  = 5 of 8
  95% CI     30.6% to 86.3%   (width 55.7 points)
  one URL    moves the figure 12.5 points

per query
  canonical tag rules                1 of 2   cited URLs ranked   (2 results captured)
  crawl budget explained             2 of 3   cited URLs ranked   (3 results captured)
  redirect chain fix                 2 of 3   cited URLs ranked   (3 results captured)

Both percentages describe these 3 queries and nothing else. Quote the figure
with its denominator and its n, or it cannot be read.
# exit code 0

Read the cited-side figures across the first three runs. One set of captures under three matching rules returns 25.0% as exact strings, 62.5% normalised and 75.0% by host: 50 points of spread from a definitional choice, on data that never changed. Every published overlap percentage rests on one of those rules and names it nowhere, and the interval beside each figure absorbs sampling error, not this. Handle the host run with care: counting a match when any page on the same host was ranked is a claim about domains, not documents.

The fourth run changes only --top, 10 to 3, which separates the denominators cleanly. Cited-side holds at 62.5% of 8, no match being deeper than rank 3, while ranked-side rises from 50.0% to 62.5% because the cut-off dropped two captured results nothing cited, shrinking its denominator from 10 to 8. A tighter cut-off flattering one number and not the other is what a single percentage hides. Every run also names log file analysis seo as excluded, having no captured results behind it, rather than scoring a miss.

The generator writes the five broken files the failure paths need, so those runs ship as a transcript rather than a claim. All eleven codes fire below, in order, verbatim.

$ python3 overlap.py cited.tsv
expected 2 file arguments, got 1
usage: overlap.py <cited.tsv> <ranked.tsv> [--top N] [--match RULE]
  cited.tsv   query <TAB> url
  ranked.tsv  query <TAB> rank <TAB> url
  --top N     keep ranks 1..N only; default 10
  --match     exact | normalised | host; default normalised
  --top and --match are the only options; -h and --help are not accepted
# exit code 2

$ python3 overlap.py cited.tsv ranked.tsv --help
'--help' is not an option this script accepts
usage: overlap.py <cited.tsv> <ranked.tsv> [--top N] [--match RULE]
  cited.tsv   query <TAB> url
  ranked.tsv  query <TAB> rank <TAB> url
  --top N     keep ranks 1..N only; default 10
  --match     exact | normalised | host; default normalised
  --top and --match are the only options; -h and --help are not accepted
# exit code 3

$ python3 overlap.py cited.tsv nowhere.tsv
input error: cannot open nowhere.tsv - no such readable file
# exit code 4

$ python3 overlap.py cited.tsv ranked.tsv --top 0
--top got '0', which is not positive
usage: overlap.py <cited.tsv> <ranked.tsv> [--top N] [--match RULE]
  cited.tsv   query <TAB> url
  ranked.tsv  query <TAB> rank <TAB> url
  --top N     keep ranks 1..N only; default 10
  --match     exact | normalised | host; default normalised
  --top and --match are the only options; -h and --help are not accepted
# exit code 5

$ python3 overlap.py cited.tsv ranked.tsv --match fuzzy
--match got 'fuzzy'; the rules are exact, normalised, host
usage: overlap.py <cited.tsv> <ranked.tsv> [--top N] [--match RULE]
  cited.tsv   query <TAB> url
  ranked.tsv  query <TAB> rank <TAB> url
  --top N     keep ranks 1..N only; default 10
  --match     exact | normalised | host; default normalised
  --top and --match are the only options; -h and --help are not accepted
# exit code 6

$ python3 overlap.py cited-header-only.tsv ranked.tsv
no data in cited-header-only.tsv: every line is blank, a comment or the header row
# exit code 7

$ python3 overlap.py cited-commas.csv ranked.tsv
wrong format in cited-commas.csv: 2 data line(s), none of them query<TAB>url - check the file is tab-separated
# exit code 8

$ python3 overlap.py cited.tsv ranked-empty.tsv
no data in ranked-empty.tsv: every line is blank, a comment or the header row
# exit code 9

$ python3 overlap.py cited.tsv ranked-headers.tsv
wrong format in ranked-headers.tsv: 2 data line(s), none of them query<TAB>rank<TAB>url - check the file is tab-separated and the rank is a number
# exit code 10

$ python3 overlap.py cited.tsv ranked-elsewhere.tsv
no query in cited.tsv appears in ranked-elsewhere.tsv at --top 10, so there is nothing to divide by
# exit code 11

$ python3 overlap.py cited-malformed.tsv ranked.tsv

matching rule      normalised
rank cut-off       top 10
queries in cited   3
queries judged     3  (a cited query with no captured results is excluded)
lines dropped      3 in cited, 0 in ranked: unparseable, and in no denominator below

CITED-SIDE
  question   how much of what was cited was also ranked
  figure     100.0%  = 3 of 3
  95% CI     43.8% to 100.0%   (width 56.2 points)
  one URL    moves the figure 33.3 points

RANKED-SIDE
  question   how much of what was ranked was also cited
  figure     30.0%  = 3 of 10
  95% CI     10.8% to 60.3%   (width 49.5 points)
  one URL    moves the figure 10.0 points

per query
  canonical tag rules                1 of 1   cited URLs ranked   (2 results captured)
  crawl budget explained             1 of 1   cited URLs ranked   (5 results captured)
  redirect chain fix                 1 of 1   cited URLs ranked   (3 results captured)

Both percentages describe these 3 queries and nothing else. Quote the figure
with its denominator and its n, or it cannot be read.
# exit code 0

The last run is the one to study. Three malformed lines get dropped, the cited-side denominator falls to 3, out comes 100.0% on 3 of 3, interval 56.2 points wide. A script swallowing those lines quietly would print the same 100.0% and nothing to explain it. Running this over a whole inventory is a scripting job, covered in Claude Code for technical SEO.

Six URL forms, three pages: a first-party measurement of the matching problem

The rule deciding when two URLs count as one is no technicality, and the cheapest demonstration is your own site. Six forms of three doctor-seo.net pages were requested on 30 September 2026, one GET each, no redirects followed: four at the Level 8 hub, one each at two pages two characters apart.

#!/usr/bin/env python3
# url-forms.py - how many URL strings reach one page, and what each one answers.
# Standard library only. One GET per form, no redirects followed.
# Run it against your own host before you trust any overlap number that
# matches URLs as strings.

import re
import sys
import urllib.error
import urllib.request

AGENT = "Mozilla/5.0 (compatible; doctor-seo.net first-party URL-form check)"


class NoRedirect(urllib.request.HTTPRedirectHandler):
    def redirect_request(self, req, fp, code, msg, headers, newurl):
        return None


def probe(url, opener):
    try:
        resp = opener.open(urllib.request.Request(url, headers={"User-Agent": AGENT}),
                           timeout=20)
    except urllib.error.HTTPError as err:
        return err.code, err.headers.get("Location"), None, None
    body = resp.read().decode("utf-8", "replace")
    resp.close()
    title = re.search(r"(?is)<title>(.*?)</title>", body)
    tag = re.search(r"(?is)<link[^>]+rel=['\"]?canonical['\"]?[^>]*>", body)
    href = re.search(r"""(?is)href=['"]?([^'"\s>]+)""", tag.group(0)) if tag else None
    return (resp.status, resp.headers.get("Location"),
            title.group(1).strip() if title else None,
            href.group(1) if href else None)


def main():
    if len(sys.argv) < 2:
        sys.stderr.write("usage: url-forms.py <url> [<url> ...]\n")
        sys.exit(2)
    opener = urllib.request.build_opener(NoRedirect)
    for url in sys.argv[1:]:
        status, location, title, canonical = probe(url, opener)
        print("%s  %s" % (status, url))
        if location:
            print("      Location:  %s" % location)
        if title:
            print("      title:     %s" % title)
        if canonical:
            print("      canonical: %s" % canonical)
        elif status == 200:
            print("      canonical: none in the served body")
    sys.exit(0)


if __name__ == "__main__":
    main()

The run, verbatim, on 30 September 2026:

$ python3 url-forms.py \
    'https://doctor-seo.net/en/seo-course/ai-for-seo/' \
    'https://doctor-seo.net/en/seo-course/ai-for-seo' \
    'http://doctor-seo.net/en/seo-course/ai-for-seo/' \
    'https://doctor-seo.net/en/seo-course/ai-for-seo/?utm_source=newsletter' \
    'https://doctor-seo.net/en/seo-course/advanced-seo-strategy/' \
    'https://doctor-seo.net/en/seo-course/advanced-seo-strategies/'

200  https://doctor-seo.net/en/seo-course/ai-for-seo/
      title:     AI for SEO: Using the Models as Tools - Doctor SEO
      canonical: https://doctor-seo.net/en/seo-course/ai-for-seo/
301  https://doctor-seo.net/en/seo-course/ai-for-seo
      Location:  https://doctor-seo.net/en/seo-course/ai-for-seo/
301  http://doctor-seo.net/en/seo-course/ai-for-seo/
      Location:  https://doctor-seo.net/en/seo-course/ai-for-seo/
200  https://doctor-seo.net/en/seo-course/ai-for-seo/?utm_source=newsletter
      title:     AI for SEO: Using the Models as Tools - Doctor SEO
      canonical: https://doctor-seo.net/en/seo-course/ai-for-seo/
200  https://doctor-seo.net/en/seo-course/advanced-seo-strategy/
      title:     Advanced SEO Strategy - Doctor SEO
      canonical: https://doctor-seo.net/en/seo-course/advanced-seo-strategy/
200  https://doctor-seo.net/en/seo-course/advanced-seo-strategies/
      title:     Advanced SEO strategies - Doctor SEO
      canonical: https://doctor-seo.net/en/seo-course/advanced-seo-strategies/
# exit code 0

Read the first four responses together. The slashed form answers 200 and declares itself canonical; the unslashed and http forms answer 301; ?utm_source=newsletter answers 200, serves the hub and names the clean slashed URL as its canonical. Four strings, one page. Exact-string matching treats them as four documents, so an assistant citing the slashed form while your capture recorded the unslashed one scores a miss on a page that was ranked. Normalised matching resolves all four to one key.

The last two responses are the warning. /en/seo-course/advanced-seo-strategy/ and /en/seo-course/advanced-seo-strategies/ both answer 200, carry different titles and each declare themselves canonical: two live pages two characters apart. Normalised matching keeps them apart; host matching merges them with everything else on the domain, which is how a 75% figure comes out of data supporting 62.5%.

One further observation explains why the script prints a title and not only a status code. Across repeated attempts on these six forms on 30 September 2026, some requests returned HTTP 200 carrying a challenge interstitial titled “One moment, please…” rather than the page. A loop recording status codes alone counts those as pages. Six forms across three attempts is a very small sample and all this site has measured; it establishes that a 200 is not proof you fetched what you asked for.

How Claude reaches a web page, and what a subprocessor listing does not tell you

Claude’s web search does not run against Google’s index, and the evidence is a compliance document rather than a product one. Xponent21 recorded Anthropic’s subprocessor listing as naming Brave Search from 19 March 2025 and turbopuffer from 6 May 2026, re-checking it on 2 July 2026: three dates doing three jobs, two subprocessor start dates and one date somebody looked. No page of that listing has been opened for this lesson, so it is attributed to Xponent21 and its re-check date.

That listing exists to satisfy data-protection obligations, worth knowing before anything is built on it. Its job is to say which outside parties touch customer data and from when, and it does that job. Questions it was never written to answer include how much of an answer each provider accounts for, whether anything reorders results afterwards, and how long the arrangement holds. The honest summary is thin: two named providers, no account of how they work together, and any step-by-step story about how your page becomes a sentence in a Claude answer is somebody’s reconstruction.

Arriving in that index at all is a separate problem with its own constraints, covered in being findable by Brave’s Web Discovery Project. Overlap picks the story up later, where something already stored gets chosen for an answer, so no amount of it substitutes for the earlier step.

One confusion sends people to the wrong log file. Anthropic documents three separate agents fetching pages from the web, ClaudeBot, Claude-User and Claude-SearchBot, each requiring its own User-agent: block, in its support article on crawling and blocking. Those agents in your logs describe Anthropic’s own fetching, not whether Brave stored your page. That page also has a documentation problem worth naming: it renders a relative timestamp and no absolute last-updated date, per this course’s live check on 23 September 2026, so a site owner acting on its agent list cannot tell how current it is.

Where Level 8 stops, Level 4 starts, and what is unsettled

Two of the three open arguments below turn on where this lesson stops. Everything above is instrumentation: capture, match, count, report an interval. The sentence that turns a figure into a recommendation about page content has crossed into Level 4, which works that for every engine rather than one. The split is how the course is organised and asserts nothing about whether the two are separate trades.

Overlap figures get carried into the argument about whether generative engine optimisation is a separate trade, and they are no use there. Read one way, 13 matches out of 15 says citations follow rankings, so the existing job covers it; read the other, it says a second index now picks who gets quoted, which is work nobody was doing. The same number supports both, recording where cited pages sat and never touching what put them there. Ending the argument would take somebody moving the same pages’ rankings and citation rates together and publishing both, and nobody has.

Markup is the next thing asked about, and the study usually produced in answer cannot settle the Claude version. Ahrefs put 1,885 pages that added JSON-LD between August 2025 and March 2026 against 4,000 that did not: AI Overviews fell 4.6%, significantly, AI Mode rose 2.4% and ChatGPT 2.2%, neither significantly. Claude is not one of those surfaces. Ahrefs adds that its pages were already cited before the markup went on, leaving the case anyone asks about, a page with no visibility yet, unmeasured. Sellers of markup as a citation lever have run past that; so have the people calling it useless.

Third argument, governing how every figure above may be reused: does a number taken inside one assistant tell you anything about a different one? Some practitioners hold the answer machinery alike enough for direction to carry; others that an index, an interface and a query mix are particular enough that a figure belongs to the product it came out of. Neither position has been tested, nobody having run two assistants side by side under one method and published it. The response here is mechanical: platform and index next to every percentage, no reading across.

Common mistakes

  • Quoting 86.7% and stopping there. Thirteen matches out of fifteen is what the figure is made of. The fix: the sample size costs four words, and those four words are what lets the sentence survive somebody checking it.
  • Comparing Profound’s figure with MERJ’s. One counts queries and the other cited URLs, so they share no scale even before the missing sample size. The fix: name the denominator in the same sentence as the number, and never average two figures whose units differ.
  • Writing an overlap figure without its engine. MERJ’s 79.2% belongs to Brave’s top 10 and its 34% to Google’s. The fix: engine, index depth and denominator, or no number.
  • Matching URLs as strings and calling it overlap. On this site four strings reach one page and two strings two characters apart reach different pages, measured 30 September 2026. The fix: choose the rule before counting, print it beside the result, and run all three rules once.
  • Asking the assistant for the number. A model with no measurement in front of it returns a plausible percentage, which is what ContextBolt’s week-long trial recorded: third-party estimates returned as factual data (ContextBolt, 23 June 2026, one site, no sample size). The fix: capture the files and do the arithmetic.

The short version

  • Profound measured 86.7% overlap between Claude’s citations and Brave’s top non-sponsored results on 21 March 2025, p<0.0001, on 15 queries. Our own division of that: 13 of 15, interval 62.1% to 96.3%.
  • The same study reports 20% for ChatGPT. At face value that is the overlap between Claude’s citations and ChatGPT’s, no comparison set recorded, so it cannot stand beside the 86.7% as one method on two engines.
  • MERJ reported in 2026 that 79.2% of Claude-cited URLs sat in Brave’s top 10 against 34% in Google’s, with no sample size, so neither carries an interval.
  • Profound counts queries, MERJ counts cited URLs: different denominators, so the two are not comparable.
  • The matching rule decides the answer: identical data returns 25.0% as exact strings, 62.5% normalised, 75.0% by host, and no published study names its rule.
  • Six URL forms over three doctor-seo.net pages, measured 30 September 2026: four strings reach the Level 8 hub, two by 301, while two strings differing by two characters are distinct self-canonical pages.
  • Xponent21 recorded Anthropic’s subprocessor listing as naming Brave Search from 19 March 2025 and turbopuffer from 6 May 2026, re-checked 2 July 2026. Such a listing names parties and dates, not a retrieval algorithm.
  • Only one claim here needs no interpretation: retrieval chooses among documents already stored, so presence is the floor rather than a dial.

Frequently asked questions

Claude cited a competitor and not me. What can I learn from that?

Very little, which is why the procedure here starts from a query list, not an incident. Capture twenty or thirty queries where you would expect to appear, record every URL cited for each, capture one engine’s results for the same queries, run the comparison. Back comes a figure with its n and a breakdown separating pages ranked but not cited from pages that were neither. Those failures need different work, and one answer cannot tell them apart.

How many queries before my overlap figure is worth quoting?

Decide from the interval, which the script prints. Eight judged URLs gave an interval over 50 points wide above; 15 queries at 86.7% gives 34.1. If your plan holds at both ends, the sample is large enough. If it changes across them, the number is not yet evidence for it.

Does a high overlap figure mean Brave rankings cause Claude’s citations?

No, and nothing published supports the causal reading. Every figure here measures agreement between two lists at one moment. Nobody has published how citations are selected or re-ranked, and no experiment has moved a ranking in one engine and measured the effect on an assistant’s citations. What is left standing is weaker and more usable: retrieval only chooses among documents already in the index, so being in it is the floor and not a dial. Monitoring your own citations and rankings together is Level 5 on SEO analytics.

I got 100% overlap on five queries. Is that a finding?

A real measurement, not yet a finding. Five out of five gives a 95% interval of roughly 57% to 100%, consistent with a true rate anywhere from a coin flip upwards. Five is also where query selection does most damage: questions chosen because you expected to appear do not behave like questions a reader asks. Report it as five of five and watch the interval close.

Sources

Entries with no link have no page in this course’s source record, checked 23 September 2026.

  • Profound, What is Claude web search, 21 March 2025 — 86.7% overlap, Claude’s citations against Brave’s top non-sponsored results, p<0.0001, n=15 queries; also 20% for ChatGPT, comparison set not stated. No link.
  • MERJ, 2026, no month published — 79.2% of Claude-cited URLs in Brave’s top 10, 34% in Google’s. No sample size published. No link.
  • A widely circulated dataset of Claude-cited URLs by domain and path — data and collection dates untraceable, so none of its figures appear here. No link.
  • Xponent21, on Anthropic’s subprocessor listing — Brave Search from 19 March 2025, turbopuffer from 6 May 2026, re-checked 2 July 2026. No link.
  • Anthropic, “Does Anthropic crawl data from the web, and how can site owners block the crawler?” — ClaudeBot, Claude-User and Claude-SearchBot, each with its own User-agent: block. Relative timestamp only, no absolute date, per live check on 23 September 2026.
  • Ahrefs, quasi-experiment on JSON-LD — 1,885 pages adding schema, August 2025 to March 2026, against 4,000 controls; AI Overviews −4.6% significant, AI Mode +2.4% and ChatGPT +2.2% not; pages already cited, Claude not measured. No link.
  • ContextBolt, “Claude SEO Experiment”, 23 June 2026 — a week of live SEO work through Claude; third-party estimates returned as factual data. One site, one team, author given only as “David”, no sample size.
  • doctor-seo.net, first-party measurement of six URL forms with url-forms.py, 30 September 2026; output above, plus two 200 responses carrying a challenge interstitial earlier that day.
  • doctor-seo.net, overlap.py, make-overlap-fixtures.sh and every run above, executed 30 September 2026: standard library only, no network access, eleven exit codes.
  • Own calculation — every 95% interval here, and the 13-of-15 derivation, comes from that script’s wilson function.

Continue the course