Skip to content
Doctor SEO

AI for SEO: Using the Models as Tools

Being Findable by Brave’s Web Discovery Project

Brave's Web Discovery Project client discards URLs on seven documented filters before it sends anything: HTTP 200 only, no redirects, no JavaScript, HTML directives only, 10 seconds, 2 MB, 22 characters of query string. Run the check on your own URLs.

Updated 30 September 2026 32 min read

By the end of this lesson you will be able to take any URL on your site and say whether Brave’s Web Discovery Project client would keep it or throw it away, before you touch a template. The check runs on two files you capture yourself: the response headers and the response body.

The finding that catches most technical SEO teams off guard: on MERJ’s reading of the Web Discovery Project client at commit f25eb3d, June 2026, that client never executes JavaScript and reads noindex and canonical only from the HTML. Note which way that cuts. The failure mode is inclusion, not exclusion: a page you took off the board with an X-Robots-Tag header is still on the board here.

Every threshold here traces to one write-up, which belongs in the opening rather than a footnote. MERJ read one open-source client at one commit and published what the code does: one analyst, one artefact, one date, no sample size and no independent replication. That establishes what the commit does. What Brave runs in production today is a separate question with no published answer.

The lesson sits in Level 8, AI for SEO, on using the models as working tools, while how AI search engines decide what to cite is Level 4, GEO and AIO. This page stays its own side of that line: a technical SEO constraint, changing what you build rather than what you write.

What you’ll learn

  • Which two of the seven filters break assumptions your stack rests on.
  • What a source-code reading of one commit does and does not establish.
  • The two routes by which a URL reaches that client: a twenty-visitor threshold, and a cheaper shortcut through search results.
  • An executable check, nine rows against one URL, with eight exit codes and the fixtures to reproduce every figure quoted.
  • Why a URL can clear all seven of MERJ’s rows and still be worthless.

The two divergences that cost the most: JavaScript and X-Robots-Tag

Two of the seven filters MERJ reported from the Web Discovery Project client at commit f25eb3d, June 2026, cut against assumptions most technical SEO work rests on: the client runs no JavaScript, and takes robots and canonical directives from the HTML alone. The other five are thresholds you can raise a ticket for. These two are architectural, so they come first.

Take JavaScript first. Googlebot has a rendering stage, queued and unpredictable but real, and a decade of front-end decisions assume something similar downstream everywhere. MERJ’s reading finds no JavaScript engine invoked anywhere in this client, so a single-page application hands Brave’s discovery path what the server sent: a mount point, a bundle tag, no sentences. Nothing can be switched on, because there is nothing to switch on.

What makes that expensive is how quietly it fails. The site still ranks in Google, coverage looks healthy in Search Console, and the route by which Claude could have reached the page was never open. No dashboard goes red. The remedy is server-side rendering or static generation, so a curl contains the sentence a reader sees. Generating pages with a model does not get you out of it: Will Scott reported in Search Engine Land on 28 August 2026 that JavaScript-heavy generated pages failed client-side rendering, in observed cases with no sample size and no control published. That article has no confirmed URL in this course’s source record, so it is cited by author and date without a link.

The second divergence inverts the usual risk. Response headers are the only place to hang a robots directive on a file with no HTML to hold one, which is why X-Robots-Tag exists and why Googlebot obeys it. On MERJ’s reading this client looks at <meta name='robots'> and nothing else, and the same goes for a canonical in a Link header. A page excluded by header alone is therefore still eligible. If a section must stay out of third-party indexes, a header is a configuration, not a guarantee: put the directive in the document too, and keep the header for the consumers that read it.

One commit, one analyst, no sample size: what this evidence is

Brave Search builds part of its index out of its own users. The Web Discovery Project is the opt-in component that reports which pages participants visit, so URLs arrive by browsing rather than by a spider walking links out of a seed set. It ships as open source, so its behaviour can be read off the code rather than guessed at from logs.

One analyst has done that reading in public. MERJ took the client at commit f25eb3d and wrote up what it does, in June 2026, and that write-up is the entire evidence base for this page.

There is no sample size to quote, which is a property of the method, not a hole in it. Reading code is not sampling behaviour: it pins down one artefact in one state, reproducible by anyone who opens the same commit and worthless beyond it. Worth saying, because the habit here is to pass these numbers around as though somebody had measured Brave’s live index. Nobody has.

What it does establish: at that commit the client checks the response status, refuses redirects, invokes no JavaScript engine, reads directives only from the HTML, and cuts off at ten seconds, two megabytes and a twenty-two character query string. Anyone with the commit hash can confirm that. What it does not establish is a longer list: which build participants run now, whether Brave applies the same criteria further down its pipeline, what proportion of real URLs fails each filter, and whether clearing one changes anything measurable. No measurement has been published pairing this client’s documented behaviour against observed indexing outcomes in Brave, and that gap is the biggest thing missing from the topic. A code reading dates fast, too: this page was written on 29 September 2026 against a June 2026 commit, and the commit is what to re-read before an expensive decision rests on it.

One neighbouring question the code reading never touches is robots.txt, and behind it, whether to block AI crawlers. That split is real: one camp argues that allowing answering agents is the only way to appear in them, the other that supplying answers without receiving visits is a transfer of value to negotiate or refuse. Neither has published a measurement that decides it, and this course does not rule on it. It is also a different mechanism, about crawlers you identify by user agent rather than discovery driven by other people’s browsing.

Why Brave’s index is the gate on whether Claude can cite you

Claude reaches the live web through Brave Search, and that dependency carries a date: Anthropic’s Trust Center lists Brave Search as a subprocessor from 19 March 2025, alongside turbopuffer from 6 May 2026. Xponent21 recorded that listing and re-checked it on 2 July 2026. It has no confirmed URL in this course’s source record, so it ships named and dated without a link.

Nothing downstream repairs a missing URL. Claude’s web search retrieves from Brave’s index, so a page Brave never stored is a page Claude cannot quote, however well written. That is why a filter list read out of a client’s source is worth the trouble: it is the part of the chain you can act on. How closely Claude’s citations track Brave’s rankings is a separate question with its own sample sizes, and it belongs to the citation lesson later in this level; no overlap percentage appears here, deliberately.

One check people run by mistake is their own log file. ClaudeBot, Claude-User and Claude-SearchBot hitting your server tells you about Anthropic’s fetching and nothing about whether Brave has stored the page. Those three agents are documented in Anthropic’s support article on its crawlers, a page that renders a relative timestamp only and so carries no date to quote.

Two ways a URL reaches the Web Discovery Project: twenty visitors, or one search result

On MERJ’s reading of commit f25eb3d, June 2026, a URL arrives at the Web Discovery Project by two routes, and they are not remotely equal in what they cost a publisher. The first is plain traffic: around twenty opted-in users, on different networks, visiting the same page, under an encryption scheme MERJ names as STAR with a twenty-share threshold. On a site that is not already large, that is not going to happen.

The second route is cheap, and it is the one to understand. One opted-in user has to see your page in a Google, Bing, Yahoo or DuckDuckGo result. The client then fetches that URL itself, after a random delay of one to twenty minutes, and it is that fetch, not the browsing, that the eligibility filters judge.

Which puts the answer where nobody selling Brave tactics wants it. Conventional rankings are the on-ramp to Brave’s index, and no Brave-specific optimisation gets round them; the groundwork is Level 2, technical SEO. Yours to control is the list of reasons that fetch gets thrown away, and the commonest is a trailing-slash redirect, of the kind covered in the lesson on redirect chains. It is the one piece of hygiene that costs you at this gate.

The eligibility check: nine rows, and the seven that belong to MERJ

The table below is the diagnostic for Brave’s Web Discovery Project client, not an illustration of one: every row is a test the script runs. Rows marked MERJ carry thresholds from that reading of commit f25eb3d, June 2026; rows marked OURS do not, and print separately.

Row Source What the client requires What the script reads Verdict on failure
200, no redirects MERJ HTTP 200, no hop Every status line DROP
At most 2 MB MERJ 2 MB limit Body size in bytes DROP
At most 10 s MERJ 10 s timeout --seconds, from %{time_total} DROP
Query string ≤ 22 MERJ Over 22 characters, rejected without a same-host canonical The query, plus the HTML <link> DROP
No noindex in HTML MERJ noindex from the HTML only Every <meta name='robots'> DROP
X-Robots-Tag agrees MERJ Robots headers are not read X-Robots-Tag against the HTML meta DIVERGE
Canonical in HTML MERJ canonical from the HTML only Link header against the HTML <link> DIVERGE
Content-Type is HTML OURS No MERJ threshold Content-Type on the final hop DROP
Text without JavaScript OURS No MERJ threshold Body text, tags and <title> stripped, against a 500 character floor DROP

A row prints one of five words. PASS and DROP read as they look. DIVERGE means the URL stays eligible but a directive you set goes unread. NOTE means there was nothing to compare, which is not passing: a response with no X-Robots-Tag gets NOTE on that row. N/A means the client never requested the body.

Two counts are easy to conflate. Nine rows, seven of them carrying MERJ’s thresholds, and those seven do not map one to one onto MERJ’s seven filters: status code and redirects share a row, HTML-directives-only splits across three, and the JavaScript filter has no row at all, because never executing JavaScript describes how the client parses rather than a rule that rejects anything. Our two rows cover what MERJ does not discuss: the served content type, and a 500 character floor on body text standing in for that missing rule.

Capture the inputs with curl, keeping every hop and the timing:

curl -sSL -D head.txt -o body.html -w '%{time_total}\n' \
  'https://example.com/technical/crawl-budget/'

Then run the check:

python3 wdp-eligibility.py 'https://example.com/technical/crawl-budget/' \
  head.txt body.html --seconds 1.18

Eight exit codes, each one failure class, none shared with an eligibility finding:

Exit Meaning
0 Kept: every blocking row passed, no divergence.
5 Kept, with a divergence: a directive Google honours that this client never reads.
1 Dropped: a blocking row failed. A drop outranks a divergence.
2 The invocation could not be parsed: wrong positional count, or an option the script does not take. There is no --help.
6 --seconds was not a positive number, or had nothing after it.
3 A named file is missing or cannot be read.
4 The header file has no HTTP status line in it.
7 The body file is empty, is a header dump, or has no HTML tag in it.

Codes 3, 4 and 7 exist so a wrong file never comes back as DROPPED. What that buys is bounded: a header dump passed as the body, an empty file and a file with no markup are caught; a truncated download is not, and the byte count on the 2 MB row is what tells you.

#!/usr/bin/env python3
# wdp-eligibility.py - would Brave's Web Discovery Project client keep this URL,
# or throw it away before it sends anything?
#
# Rows marked MERJ carry thresholds MERJ reported from the Web Discovery Project
# client at commit f25eb3d, June 2026. Rows marked OURS are two additions by
# doctor-seo.net; MERJ publishes no threshold for either, and they print in a
# separate block so they never borrow MERJ's authority.
#
# No network access. Capture the two files yourself, then run it:
#   curl -sSL -D head.txt -o body.html -w '%{time_total}\n' 'https://example.com/p/'
#   python3 wdp-eligibility.py 'https://example.com/p/' head.txt body.html --seconds 1.52
#
# Row states: PASS, DROP (blocking), DIVERGE (eligible but a directive is unread),
# NOTE (nothing to compare against), N/A (the client never requested this body).
#
# Exit codes. Each names one failure class, and no input fault shares a code with
# a real eligibility finding, so a wrong file can never be read as a result.
#   0  KEPT     - every blocking row passed and no divergence was found
#   5  KEPT, WITH A DIVERGENCE - eligible, but a directive Google honours is one
#                this client never reads, so you are in when you meant to be out
#   1  DROPPED  - at least one blocking row failed; a drop outranks a divergence
#   2  the invocation could not be parsed: wrong positional count, or an option
#      this script does not take (it takes no -h and no --help)
#   3  a named file is missing or cannot be read
#   4  the header file is unusable: no HTTP status line in it
#   7  the body file is unusable: empty, an HTTP header dump, or no HTML tag
#   6  --seconds was not a positive number

import os
import re
import sys
from urllib.parse import urlparse

CAP_BYTES = 2 * 1024 * 1024   # MERJ: 2 MB
CAP_SECONDS = 10.0            # MERJ: 10 s
CAP_QUERY = 22                # MERJ: query strings longer than this are rejected
MIN_TEXT = 500                # OURS, not MERJ's. Tune it to your templates.

USAGE = ("usage: wdp-eligibility.py <url> <head.txt> <body.html> [--seconds N]\n"
         "  head.txt   response headers, e.g. curl -D head.txt (keep every hop)\n"
         "  body.html  response body for the SAME request\n"
         "  --seconds  total time from curl -w '%{time_total}', a positive number\n"
         "  no other options are accepted, including -h and --help")

rows = []     # each: (source, label, state, note)


def add(source, label, state, note):
    rows.append((source, label, state, note))


def die(code, message):
    sys.stderr.write(message.rstrip() + "\n")
    sys.exit(code)


def slurp(path):
    if not os.path.isfile(path):
        die(3, "input error: not a readable file: %s" % path)
    try:
        with open(path, "rb") as fh:
            return fh.read()
    except OSError as err:
        die(3, "input error: cannot read %s (%s)" % (path, err.strerror))


def parse_hops(raw):
    """Split a curl -D dump into hops. Returns [] if it holds no status line."""
    hops, current = [], None
    for line in raw.decode("utf-8", "replace").splitlines():
        if re.match(r"^HTTP/\d", line):
            current = {"status": line.strip(), "headers": {}}
            hops.append(current)
        elif current is not None and ":" in line:
            name, value = line.split(":", 1)
            current["headers"].setdefault(name.strip().lower(), []).append(value.strip())
    return hops


def status_of(hop):
    found = re.match(r"^HTTP/\S+\s+(\d{3})", hop["status"])
    return int(found.group(1)) if found else 0


def stripped_text(markup):
    without = re.sub(r"(?is)<(script|style|template|noscript|title)\b.*?</\1>", " ", markup)
    without = re.sub(r"(?s)<!--.*?-->", " ", without)
    without = re.sub(r"(?s)<[^>]+>", " ", without)
    return re.sub(r"\s+", " ", without).strip()


def html_canonical(markup):
    tag = re.search(r"(?is)<link[^>]+rel=['\"]?canonical['\"]?[^>]*>", markup)
    if not tag:
        return None
    href = re.search(r"""(?is)href=['"]?([^'"\s>]+)""", tag.group(0))
    return href.group(1) if href else None


def judge(url, head_path, body_path, seconds):
    hops = parse_hops(slurp(head_path))
    if not hops:
        die(4, "header file unusable: no HTTP status line in %s - wrong file, or curl "
               "wrote the body here" % head_path)

    raw_body = slurp(body_path)
    if not raw_body.strip():
        die(7, "body file unusable: %s is empty, so there is no document to judge" % body_path)
    body = raw_body.decode("utf-8", "replace")
    if re.match(r"^\s*HTTP/\d", body):
        die(7, "body file unusable: %s starts with an HTTP status line, so it is a header "
               "dump and not a response body" % body_path)
    if "<" not in body:
        die(7, "body file unusable: no HTML tag anywhere in %s" % body_path)

    size = len(raw_body)
    hop_codes = [status_of(hop) for hop in hops]
    redirects = [code for code in hop_codes if 300 <= code < 400]
    last = hops[-1]["headers"]

    # Row 1 (MERJ): HTTP 200 only, and no redirect anywhere in the chain.
    if redirects:
        add("MERJ", "200, no redirects", "DROP",
            "first hop is %d, and the client drops the URL at the hop instead of "
            "following it" % redirects[0])
        dropped_early = True
    elif hop_codes[-1] != 200:
        add("MERJ", "200, no redirects", "DROP",
            "final response %d; only 200 is accepted" % hop_codes[-1])
        dropped_early = True
    else:
        add("MERJ", "200, no redirects", "PASS", "single 200, no intermediate hop")
        dropped_early = False

    def body_row(source, label, state, note):
        # After a drop at the hop the client never requests this body, so nothing
        # measured from it describes the URL under test.
        if dropped_early:
            add(source, label, "N/A", "not applicable: the client never requests this body")
        else:
            add(source, label, state, note)

    # Row 2 (MERJ): 2 MB cap.
    body_row("MERJ", "at most 2 MB", "PASS" if size <= CAP_BYTES else "DROP",
             "%d bytes, %.1f%% of the cap" % (size, size / CAP_BYTES * 100))

    # Row 3 (MERJ): 10 second timeout.
    if seconds is None:
        body_row("MERJ", "at most 10 s", "NOTE",
                 "not measured: pass --seconds with curl's %{time_total}")
    else:
        body_row("MERJ", "at most 10 s", "PASS" if seconds <= CAP_SECONDS else "DROP",
                 "%.2f s against a %.0f s ceiling" % (seconds, CAP_SECONDS))

    # Row 4 (MERJ): query string over 22 chars needs a same-host canonical in the HTML.
    query = urlparse(url).query
    canonical = html_canonical(body)
    same_host = bool(canonical) and urlparse(canonical).netloc in ("", urlparse(url).netloc)
    if len(query) <= CAP_QUERY:
        body_row("MERJ", "query string <= 22", "PASS",
                 "%d characters, under the limit" % len(query))
    elif same_host:
        body_row("MERJ", "query string <= 22", "PASS",
                 "%d characters, rescued by a same-host canonical: %s" % (len(query), canonical))
    else:
        body_row("MERJ", "query string <= 22", "DROP",
                 "%d characters and no same-host canonical in the HTML" % len(query))

    # Row 5 (MERJ): noindex read from the HTML only.
    metas = re.findall(r"(?is)<meta[^>]+name=['\"]?robots['\"]?[^>]*>", body)
    meta_text = " ".join(metas).lower()
    body_row("MERJ", "no noindex in HTML", "PASS" if "noindex" not in meta_text else "DROP",
             metas[0][:96] if metas else "no meta robots present")

    # Row 6 (MERJ): a robots directive sent as a header is never read.
    header_robots = last.get("x-robots-tag") or []
    if not header_robots:
        body_row("MERJ", "X-Robots-Tag agrees", "NOTE", "no X-Robots-Tag header sent")
    elif "noindex" in " ".join(header_robots).lower() and "noindex" not in meta_text:
        body_row("MERJ", "X-Robots-Tag agrees", "DIVERGE",
                 "noindex sits in the header only (%s): Google honours it, this client "
                 "does not read it" % header_robots[0])
    else:
        body_row("MERJ", "X-Robots-Tag agrees", "PASS",
                 "header present and consistent with the HTML: %s" % header_robots[0])

    # Row 7 (MERJ): a canonical sent as a header is never read either.
    header_link = [v for v in (last.get("link") or []) if "canonical" in v.lower()]
    header_canonical = None
    if header_link:
        found = re.search(r"<([^>]+)>", header_link[0])
        header_canonical = found.group(1).strip() if found else header_link[0]
    if header_canonical and not canonical:
        body_row("MERJ", "canonical in HTML", "DIVERGE",
                 "canonical sits in the Link header only (%s); this client reads only the "
                 "HTML <link>" % header_canonical)
    elif header_canonical and canonical and header_canonical.rstrip("/") != canonical.rstrip("/"):
        body_row("MERJ", "canonical in HTML", "DIVERGE",
                 "the Link header says %s and the HTML says %s; this client obeys the HTML "
                 "and ignores the header" % (header_canonical, canonical))
    elif canonical:
        body_row("MERJ", "canonical in HTML", "PASS", canonical)
    else:
        body_row("MERJ", "canonical in HTML", "NOTE", "no canonical declared anywhere")

    # ---- OURS from here: doctor-seo.net's two additions, not MERJ's findings ----
    content_type = (last.get("content-type") or [""])[0]
    body_row("OURS", "Content-Type is HTML", "PASS" if "html" in content_type.lower() else "DROP",
             content_type or "no Content-Type header")

    shell = re.search(r"(?is)<div[^>]+id=['\"]?(root|app|__next|__nuxt)['\"]?[^>]*>\s*</div>", body)
    text = stripped_text(body)
    if shell:
        body_row("OURS", "text without JavaScript", "DROP",
                 "empty application shell: %s" % shell.group(0)[:64])
    else:
        body_row("OURS", "text without JavaScript",
                 "PASS" if len(text) >= MIN_TEXT else "DROP",
                 "%d characters of body text (our floor is %d)" % (len(text), MIN_TEXT))


def report(source, heading):
    print(heading)
    for src, label, state, note in rows:
        if src == source:
            print("  [%-7s] %-24s %s" % (state, label, note))
    print()


def main():
    argv = sys.argv[1:]
    seconds = None
    if "--seconds" in argv:
        at = argv.index("--seconds")
        if at + 1 >= len(argv):
            die(6, "--seconds needs a positive number after it\n" + USAGE)
        try:
            seconds = float(argv[at + 1])
        except ValueError:
            die(6, "--seconds got %r, which is not a number\n%s" % (argv[at + 1], USAGE))
        if seconds <= 0:
            die(6, "--seconds got %r, which is not positive\n%s" % (argv[at + 1], USAGE))
        argv = argv[:at] + argv[at + 2:]
    unknown = [a for a in argv if a.startswith("-")]
    if unknown:
        die(2, "invocation error: this script takes no option %r\n%s" % (unknown[0], USAGE))
    if len(argv) != 3:
        die(2, "invocation error: expected 3 arguments, got %d\n%s" % (len(argv), USAGE))

    judge(argv[0], argv[1], argv[2], seconds)

    print("\nURL: %s\n" % argv[0])
    report("MERJ", "MERJ's rows - thresholds reported from the Web Discovery Project client\n"
                   "at commit f25eb3d, June 2026.")
    report("OURS", "Our rows - doctor-seo.net's own two checks, NOT MERJ's: the served\n"
                   "content type, and %d characters of body text as a stand-in for\n"
                   "\"the client never executes JavaScript\". MERJ publishes no text floor."
                   % MIN_TEXT)

    drops = sum(1 for _, _, state, _ in rows if state == "DROP")
    diverges = sum(1 for _, _, state, _ in rows if state == "DIVERGE")
    if drops:
        print("DROPPED: %d blocking row(s) failed, so the client discards this URL." % drops)
        sys.exit(1)
    if diverges:
        print("KEPT, WITH %d DIVERGENCE(S): eligible, but a directive you set is one this\n"
              "client never reads. You are in when you meant to be out." % diverges)
        sys.exit(5)
    print("KEPT: every blocking row passed. That is eligibility, not indexing - one of the\n"
          "two discovery routes still has to happen.")
    sys.exit(0)


if __name__ == "__main__":
    main()

What it says about real URLs, and four fixtures that reproduce every figure here

Three real URLs have been through this check, measured on doctor-seo.net on 25 September 2026 and published in this lesson’s Spanish edition. The Level 8 hub answered 200 directly, 56,962 bytes, 2.7% of the two megabyte cap, in 1.52 seconds, with its canonical in the HTML. The same URL without its trailing slash returned a 301, which drops it at the hop. The course index carrying ?utm_source=boletin&utm_medium=email has a 35 character query string and survived on its same-host canonical. Three URLs on one host: a small first-party sample, and all this site has measured.

Everything below is constructed fixtures, not measurements of anything live. A fixture can be shipped, so anyone running the script against it gets the output printed here byte for byte, today or in two years. Save it as make-fixtures.sh and run it in an empty directory:

#!/bin/sh
# make-fixtures.sh - writes the five files the four worked examples use.
# These are CONSTRUCTED fixtures, not measurements of any live site.

printf 'HTTP/2 200 \r\ncontent-type: text/html; charset=utf-8\r\ncache-control: max-age=600\r\n\r\n' > f1.head.txt

cat > f1.body.html <<'HTML'
<!doctype html>
<html lang="en"><head>
<meta charset="utf-8">
<title>Crawl budget: what it is and when it matters</title>
<link rel="canonical" href="https://example.com/technical/crawl-budget/">
</head><body>
<h1>Crawl budget: what it is and when it matters</h1>
<p>Crawl budget is the number of URLs a search engine is willing to request from
your host in a given period. On a site of a few hundred pages it is not a problem
you have. On a site with faceted navigation, a calendar and a search template that
emits a unique URL for every query, it decides whether this week's new products are
found now or next quarter.</p>
<p>Two things set it: how much the engine wants your pages, and how fast your host
answers. You move the first with internal linking and by deleting the templates
that manufacture near-duplicate URLs. You move the second with caching, and by not
doing database work on a page that could have been a file.</p>
<p>The diagnostic is a log file. Group requests by user agent, then by path pattern,
and look at where the requests went that you would not have chosen to spend. A site
spending forty per cent of its crawl on parameter permutations of one listing does
not need a bigger budget. It needs fewer URLs.</p>
</body></html>
HTML

printf 'HTTP/2 301 \r\nlocation: https://example.com/technical/crawl-budget/\r\ncontent-length: 0\r\n\r\nHTTP/2 200 \r\ncontent-type: text/html; charset=utf-8\r\n\r\n' > f2.head.txt

printf 'HTTP/2 200 \r\ncontent-type: text/html; charset=utf-8\r\nx-robots-tag: noindex, nofollow\r\n\r\n' > f3.head.txt

cat > f4.body.html <<'HTML'
<!doctype html>
<html lang="en"><head><meta charset="utf-8"><title>Crawl budget</title></head>
<body><div id="root"></div><script src="/bundle.js"></script></body></html>
HTML

The four runs, verbatim, as produced on 29 September 2026:

$ python3 wdp-eligibility.py 'https://example.com/technical/crawl-budget/' \
    f1.head.txt f1.body.html --seconds 1.18

URL: https://example.com/technical/crawl-budget/

MERJ's rows - thresholds reported from the Web Discovery Project client
at commit f25eb3d, June 2026.
  [PASS   ] 200, no redirects        single 200, no intermediate hop
  [PASS   ] at most 2 MB             1252 bytes, 0.1% of the cap
  [PASS   ] at most 10 s             1.18 s against a 10 s ceiling
  [PASS   ] query string <= 22       0 characters, under the limit
  [PASS   ] no noindex in HTML       no meta robots present
  [NOTE   ] X-Robots-Tag agrees      no X-Robots-Tag header sent
  [PASS   ] canonical in HTML        https://example.com/technical/crawl-budget/

Our rows - doctor-seo.net's own two checks, NOT MERJ's: the served
content type, and 500 characters of body text as a stand-in for
"the client never executes JavaScript". MERJ publishes no text floor.
  [PASS   ] Content-Type is HTML     text/html; charset=utf-8
  [PASS   ] text without JavaScript  996 characters of body text (our floor is 500)

KEPT: every blocking row passed. That is eligibility, not indexing - one of the
two discovery routes still has to happen.
# exit code 0

$ python3 wdp-eligibility.py 'https://example.com/technical/crawl-budget' \
    f2.head.txt f1.body.html --seconds 1.18

URL: https://example.com/technical/crawl-budget

MERJ's rows - thresholds reported from the Web Discovery Project client
at commit f25eb3d, June 2026.
  [DROP   ] 200, no redirects        first hop is 301, and the client drops the URL at the hop instead of following it
  [N/A    ] at most 2 MB             not applicable: the client never requests this body
  [N/A    ] at most 10 s             not applicable: the client never requests this body
  [N/A    ] query string <= 22       not applicable: the client never requests this body
  [N/A    ] no noindex in HTML       not applicable: the client never requests this body
  [N/A    ] X-Robots-Tag agrees      not applicable: the client never requests this body
  [N/A    ] canonical in HTML        not applicable: the client never requests this body

Our rows - doctor-seo.net's own two checks, NOT MERJ's: the served
content type, and 500 characters of body text as a stand-in for
"the client never executes JavaScript". MERJ publishes no text floor.
  [N/A    ] Content-Type is HTML     not applicable: the client never requests this body
  [N/A    ] text without JavaScript  not applicable: the client never requests this body

DROPPED: 1 blocking row(s) failed, so the client discards this URL.
# exit code 1

$ python3 wdp-eligibility.py 'https://example.com/technical/crawl-budget/?utm_source=newsletter&utm_medium=email' \
    f3.head.txt f1.body.html --seconds 1.18

URL: https://example.com/technical/crawl-budget/?utm_source=newsletter&utm_medium=email

MERJ's rows - thresholds reported from the Web Discovery Project client
at commit f25eb3d, June 2026.
  [PASS   ] 200, no redirects        single 200, no intermediate hop
  [PASS   ] at most 2 MB             1252 bytes, 0.1% of the cap
  [PASS   ] at most 10 s             1.18 s against a 10 s ceiling
  [PASS   ] query string <= 22       38 characters, rescued by a same-host canonical: https://example.com/technical/crawl-budget/
  [PASS   ] no noindex in HTML       no meta robots present
  [DIVERGE] X-Robots-Tag agrees      noindex sits in the header only (noindex, nofollow): Google honours it, this client does not read it
  [PASS   ] canonical in HTML        https://example.com/technical/crawl-budget/

Our rows - doctor-seo.net's own two checks, NOT MERJ's: the served
content type, and 500 characters of body text as a stand-in for
"the client never executes JavaScript". MERJ publishes no text floor.
  [PASS   ] Content-Type is HTML     text/html; charset=utf-8
  [PASS   ] text without JavaScript  996 characters of body text (our floor is 500)

KEPT, WITH 1 DIVERGENCE(S): eligible, but a directive you set is one this
client never reads. You are in when you meant to be out.
# exit code 5

$ python3 wdp-eligibility.py 'https://example.com/app/' \
    f1.head.txt f4.body.html --seconds 1.18

URL: https://example.com/app/

MERJ's rows - thresholds reported from the Web Discovery Project client
at commit f25eb3d, June 2026.
  [PASS   ] 200, no redirects        single 200, no intermediate hop
  [PASS   ] at most 2 MB             171 bytes, 0.0% of the cap
  [PASS   ] at most 10 s             1.18 s against a 10 s ceiling
  [PASS   ] query string <= 22       0 characters, under the limit
  [PASS   ] no noindex in HTML       no meta robots present
  [NOTE   ] X-Robots-Tag agrees      no X-Robots-Tag header sent
  [NOTE   ] canonical in HTML        no canonical declared anywhere

Our rows - doctor-seo.net's own two checks, NOT MERJ's: the served
content type, and 500 characters of body text as a stand-in for
"the client never executes JavaScript". MERJ publishes no text floor.
  [PASS   ] Content-Type is HTML     text/html; charset=utf-8
  [DROP   ] text without JavaScript  empty application shell: <div id="root"></div>

DROPPED: 1 blocking row(s) failed, so the client discards this URL.
# exit code 1

Fixture one is a clean page: eight rows PASS, one prints NOTE because there is no X-Robots-Tag to compare against, exit 0. Fixture two is the same page without its trailing slash, and it is the instructive one: row one fails, the other eight print N/A rather than a number, and exit is 1. That N/A is the most useful decision in the output, because the client never asks for a body it has already dropped, so what curl saved describes the redirect target instead. Fixture three is a campaign URL whose 38 character query string passes on the HTML’s same-host canonical, demonstrating the documented exception instead of asserting it, while its X-Robots-Tag: noindex with no matching <meta name='robots'> prints DIVERGE and exits 5.

Fixture four is why the two OURS rows exist. A 171 byte document holding <div id='root'></div> and a bundle tag clears all seven of MERJ’s rows: five PASS, two NOTE, not one DROP. Only our text row objects, exiting 1. Clearing the published filters is not the same as being worth storing.

The input faults were exercised too: body file as the header dump exits 4; header dump as the body, and an empty body, exit 7; a missing filename exits 3; two positional arguments and --help exit 2; --seconds fast exits 6. Running a whole inventory through the same nine rows is a scripting job, and the Claude Code for technical SEO lesson covers that work.

Brave’s filters are a build constraint, not a visibility tactic

Everything the seven Web Discovery Project filters imply is a construction decision, which is why this lesson sits with the technical material: nothing here tells you what to write, only whether a document you already wrote can be fetched, parsed and stored. “Claude uses Brave, so rank in Brave” gets written up regularly and is less wrong than unusable: it names a dependency and stops. The layer underneath is response codes and parsing rules.

Whether generative engine optimisation is a discipline separate from search engine optimisation is genuinely disputed: one camp points at engines with their own indexes and retrieval behaviour, the other at how much of what works is crawling and indexing relabelled. This page is evidence for neither, because a client that refuses redirects and never renders is a crawling problem in either vocabulary. That argument lives in Level 4, and what a model can be handed in the first place is what Claude can and cannot do for SEO.

What to change this week

Five actions, none of them new advice. Against the Web Discovery Project client as MERJ documented it in June 2026, each moves from recommended to blocking.

  1. Serve your text in the source HTML. curl a page and search the output for a sentence you see on screen. If it is absent, this client never sees it either.
  2. Declare canonical and robots in the HTML. Keep the headers for other consumers, but never let one be the only place a directive lives.
  3. Circulate the URL that answers 200 directly. The client judges what the opted-in user saw and does not follow hops. Audit your newsletter, social profiles and ad destinations for unslashed variants.
  4. Check your parameterised URLs. Any realistic campaign string clears 22 characters, so the same-host canonical does all the work. Find the templates that omit it.
  5. Measure weight and response time cold: two megabytes and ten seconds, on an uncached request at your busiest hour.

Common mistakes

  1. Treating the seven thresholds as permanent facts. They come from one commit read in June 2026, not a published specification. The fix: store the date beside every threshold and re-read the commit before an architectural decision rests on one.
  2. Excluding content with X-Robots-Tag alone. Google honours it, this client does not read it, so you stay in an index you believed you had left. The fix: mirror the directive in <meta name='robots'> and verify it in the served HTML, not the config.
  3. Accepting a URL because it “opens fine in the browser”. A browser follows redirects and runs JavaScript; this client does neither. The fix: curl and the nine-row check.
  4. Reading NOTE as PASS. A NOTE means nothing was there to compare, so the row proves nothing. The fix: count only PASS rows when you report a result.

The short version

  • Claude reaches the web through Brave Search: Anthropic’s Trust Center lists it as a subprocessor from 19 March 2025, per Xponent21, re-checked 2 July 2026.
  • Two discovery routes: roughly twenty opted-in users on different networks visiting the page under a STAR twenty-share threshold, or one opted-in user seeing it in a Google, Bing, Yahoo or DuckDuckGo result, after which the client re-requests the URL following a random one to twenty minute delay.
  • Seven filters can discard that request: HTTP 200 only, no redirects, JavaScript never executed, noindex and canonical read only from the HTML, 10 seconds, 2 MB, and query strings over 22 characters unless a same-host canonical saves them.
  • All seven come from one source: MERJ’s reading of the client at commit f25eb3d, June 2026, with no sample size because reading code is not sampling behaviour. No measurement pairs it against observed indexing in Brave.
  • Conventional rankings are the practical on-ramp, the search-result route being the only affordable one.
  • A header-only noindex inverts the usual risk: the page stays eligible, so a header is not proof of exclusion.
  • A document can clear all seven filters and still be an empty shell: eligibility is a floor, not a result.

Frequently asked questions

If every row passes, am I in Brave’s index?

No, and the script says so on the way out. Clearing the nine rows means the Web Discovery Project client would not throw the URL away; it says nothing about whether the client ever sees it. That needs one of the two discovery routes: roughly twenty opted-in users on separate networks, or one who meets the page in a search result. Eligibility is the part you control.

Which should I fix first, the redirects or the rendering?

The redirects, because they are cheap and the rendering is not. Circulating the URL that answers 200 directly is an afternoon auditing where your links point; moving off client-side rendering has a budget attached. Run the check on twenty representative URLs and let it order the work: a DROP on row one is a link-hygiene bug, a DROP on the text row is a roadmap item.

Does this work on a competitor’s URLs?

Yes, and it is the cheapest use of it. The check needs nothing but a public response, so curl a rival’s key pages and run the nine rows. Do not over-read the output: it describes eligibility under one commit’s rules, not whether Claude cites them.

What would make this page wrong?

Two things. A new commit that changes a threshold makes every number here out of date, which is why the hash and the June 2026 date sit next to each one. The other is better: a published measurement pairing the client’s behaviour against real indexing outcomes in Brave. A URL indexed in Brave that fails one of these rows beats any restatement of a threshold, and setting that monitoring up is Level 5, SEO analytics.

Sources

Entries marked (no link) have no confirmed URL in this course’s source record, checked 23 September 2026.

  • MERJ, reading of the open-source Brave Web Discovery Project client at commit f25eb3d, June 2026. No sample size: one artefact, one commit, one analyst, no replication. (no link)
  • Anthropic Trust Center subprocessor listing: Brave Search from 19 March 2025, turbopuffer from 6 May 2026, recorded by Xponent21, re-checked 2 July 2026. (no link)
  • Anthropic, “Does Anthropic crawl data from the web, and how can site owners block the crawler?”, on ClaudeBot, Claude-User and Claude-SearchBot. The page renders a relative timestamp only, so no absolute date is quoted.
  • Will Scott, “Use Claude for SEO. Don’t let Claude do SEO.”, Search Engine Land, 28 August 2026. Observed cases, no sample size, no control. (no link)
  • doctor-seo.net, own measurement of three URLs, 25 September 2026, published in this lesson’s Spanish edition. Unlinked: this site does not link across languages.
  • wdp-eligibility.py, make-fixtures.sh and the four runs, published above and executed on 29 September 2026. Standard library only, no network, eight exit codes.

Continue the course