Skip to content
Doctor SEO

AI for SEO: Using the Models as Tools

Why Every Chatbot Invents Search Volumes, and What to Do Instead

A general-purpose chatbot with no connected data source holds no search-volume dataset, so the number it gives you was composed. Here is the four-question provenance test, and a script that runs it over a whole export.

Updated 3 October 2026 26 min read

Ask any general-purpose chatbot what a keyword gets per month and it will give you a number. Ask where the number came from and the answer thins out. That gap is the lesson: a model with nothing wired to it has no search-volume dataset to read, so the figure was composed, not looked up.

By the end you will have a provenance test for any figure a model produces. Four questions, one verdict, and a Python script that runs the same checks across a whole keyword export and will not let a damaged or empty file come back clean.

Two neighbouring jobs are out of scope. Choosing which engine to open for which task was the previous lesson’s routing table in this level. Getting an AI search engine to cite your pages belongs to the level on GEO and AIO, where whether that is a separate discipline at all is left open rather than settled.

What you’ll learn

  • Put a four-question provenance test to any figure a model gives you, by hand, in a minute.
  • Defend a priority order in a client meeting with no volume column in the sheet, naming the substitute behind it and what that substitute measures.
  • Explain why no general-purpose model holds search-volume data, and why that is structural rather than a defect a vendor will patch.
  • State what connecting a real data source fixes, and what it leaves untouched.
  • Run the same checks over a whole export with a script that fails loudly on an empty file.

The four questions to put to any figure a model gives you

The provenance test is four questions, put to the figure rather than to the model that produced it. Who measured this? When was it pulled? What was counted? Who is it about? A figure that cannot answer all four was generated, and a generated figure does not enter a deliverable. There is no partial credit.

The verdict is not true or false. It is auditable or not auditable, and its working form is a repeat test: given only your four answers, could a second person pull the number again? If they could, the figure is defensible even if it later proves a poor estimate. If not, it is text shaped like data.

Question An answer that passes An answer that fails
Who measured it? A named tool, account or body somebody can go back to The name of a model, “several sources”, silence
When was it pulled? The export date, to the day “Recently”, a bare year, nothing
What was counted? The window and arithmetic the provider documents: twelve-month average, median, modelled or observed “It is what the tool gives”, an empty field
Who is it about? Market, language and period covered A two-letter country code on its own
The verdict Auditable: a second person could repeat the pull from these answers alone Not auditable: the cell stays empty

Why the answer is the same whichever chatbot you asked

Where does a search volume actually live? Inside a provider’s product. Somebody runs a panel or a clickstream, fits a model to what it records, and sells the result by keyword, market and language to account holders. Training a language model on web text does not put that product inside it. Ask with no connector attached and back comes a completion of a sentence that ends in a number.

The reason this gets through review is that the output is well formed. Volumes have conventions: an order of magnitude, a habit of rounding, a predictable fall from head term to long tail. Conventions are what a language model reproduces well, and a conventional-looking figure is not a measured one. In a cell, nothing records which of the two you were handed.

This belongs at category level, not as a complaint about a product. No general-purpose model has a search-volume dataset behind it, which follows from asking any system for a quantity it cannot reach rather than from a bug one rival has fixed. It is why keyword research in ChatGPT and in Gemini need one answer, not two.

The quotable line on this belongs to a company that sells keyword data. Si Quan Ong put it on the Ahrefs blog, 30 October 2025: the volumes a model quotes are “simply guesses”. The piece argues a case rather than reporting one and carries no sample size, and it is named and dated here rather than linked.

The empty cell: what to hand over when there is no number

The deliverable question answers itself once the provenance test is running: ship the gap, labelled. A row reading “no volume from a nameable source, prioritised on autocomplete depth, which is ordinal” carries more than a figure nobody can defend, because it says what you did instead. The alternative costs you later, once the figure is in a forecast and somebody starts tracing it.

Dates belong in the same row as the figures that did pass, not in a footnote. A figure with its export date attached ages where you can see it; one without is read as current for as long as the file survives. Keyword research as a discipline is taught in the fundamentals level, and reading demand measured on your own property is the analytics level.

Running the four questions over a whole export

Four questions are manageable for one figure in a chat window. A keyword deliverable has hundreds of rows, and the failure that costs money is a file that looks finished. The script below puts the four questions to an export header, checks every row, and is built so a file saying nothing cannot pass.

The exit codes are numbered around a hazard. Python exits 1 on an uncaught exception, and a wrapper round main() cannot catch everything: a write to a full or broken stdout fails inside the printing itself. So 1 is left unassigned, and seeing it means the run broke rather than that any verdict was reached. A file that failed its checks and a file nobody could parse get separate numbers too.

Exit code Verdict What produced it
0 AUDITABLE Header complete, every row usable, nothing flagged
1 unassigned Reserved: a failure this script cannot catch
2 AUDITABLE WITH WARNINGS Nothing failed; something is worth a look
3 NOT AUDITABLE A check failed, including an export with no usable value
4 INPUT UNUSABLE No argument given, missing file, unreadable file, invalid JSON, unexpected top level
5 INTERNAL ERROR An exception the wrapper caught
#!/usr/bin/env python3
"""volume_provenance.py - audit the provenance of a keyword export.

Usage:
    python3 volume_provenance.py export.json
    python3 volume_provenance.py --selftest      # deliberate fault injection

The export is one JSON object: an "export" header declaring where the
numbers came from, and a "rows" list of {"keyword": ..., "volume": ...}.

Exit codes. 1 is deliberately left unassigned, because Python itself
exits 1 on an uncaught exception and on a failure this script cannot
catch, such as a write to a full or broken stdout. Seeing 1 therefore
means the run broke, never that your export was judged:
    0  AUDITABLE - header complete, every row usable, no warnings
    1  (unassigned - a failure outside this script's control)
    2  AUDITABLE WITH WARNINGS - nothing failed, something is worth a look
    3  NOT AUDITABLE - at least one check failed
    4  INPUT UNUSABLE - no argument, missing file, bad JSON, wrong shape
    5  INTERNAL ERROR - an exception the wrapper caught; not about you

Standard library only. Makes no network calls. Checks whether a number
can be traced, never whether it is correct.
"""

import json
import re
import sys
from datetime import date, datetime

NOT_A_SOURCE = ("chatgpt", "gpt-", "gemini", "claude", "perplexity", "copilot",
                "grok", "deepseek", "the model", "the assistant", "the ai",
                "an ai", "ai estimate", "various sources", "industry data",
                "general knowledge", "best guess")

EVASIVE_METHOD = ("not stated", "unknown", "no idea", "not sure", "whatever",
                  "the tool just gives it", "the model worked it out")

ISO_DAY = re.compile(r"^\d{4}-\d{2}-\d{2}$")

STALE_DAYS = 180
MIN_SOURCE = 6
MIN_METHOD = 20
MIN_SCOPE = 8


class Report:
    def __init__(self):
        self.lines = []
        self.fails = 0
        self.warns = 0

    def ok(self, label, detail):
        self.lines.append("  [ OK   ] %-11s %s" % (label, detail))

    def warn(self, label, detail):
        self.warns += 1
        self.lines.append("  [ WARN ] %-11s %s" % (label, detail))

    def fail(self, label, detail):
        self.fails += 1
        self.lines.append("  [ FAIL ] %-11s %s" % (label, detail))

    def note(self, text):
        self.lines.append(text)

    def dump(self):
        for line in self.lines:
            print(line)


def check_source(value, rep):
    text = value.strip() if isinstance(value, str) else ""
    if len(text) < MIN_SOURCE:
        rep.fail("source", "missing or too short to identify: %r" % (value,))
        return
    low = text.lower()
    for token in NOT_A_SOURCE:
        if token in low:
            rep.fail("source", "names a generator, not a measurer: %r" % text)
            return
    rep.ok("source", text)


def check_retrieved(value, rep):
    text = value.strip() if isinstance(value, str) else ""
    if not ISO_DAY.match(text):
        rep.fail("retrieved", "not an ISO date (YYYY-MM-DD): %r" % (value,))
        return
    try:
        pulled = datetime.strptime(text, "%Y-%m-%d").date()
    except ValueError:
        rep.fail("retrieved", "not a real calendar date: %r" % text)
        return
    today = date.today()
    if pulled > today:
        rep.fail("retrieved", "dated in the future: %s" % text)
        return
    if (today - pulled).days > STALE_DAYS:
        rep.warn("retrieved", "%s, over %d days old" % (text, STALE_DAYS))
    else:
        rep.ok("retrieved", text)


def check_method(value, rep):
    text = value.strip() if isinstance(value, str) else ""
    if len(text) < MIN_METHOD:
        rep.fail("method", "no collection method described: %r" % (value,))
        return
    low = text.lower()
    for token in EVASIVE_METHOD:
        if token in low:
            rep.fail("method", "declared unknown: %r" % text)
            return
    rep.ok("method", text)


def check_scope(value, rep):
    text = value.strip() if isinstance(value, str) else ""
    if len(text) < MIN_SCOPE:
        rep.fail("scope", "no market, language or period given: %r" % (value,))
        return
    parts = [p for p in text.split(",") if p.strip()]
    if len(parts) < 2:
        rep.fail("scope", "one fragment only, needs market and period: %r" % text)
        return
    rep.ok("scope", text)


def usable_volume(value):
    """Return the number, or None. bool is rejected: it subclasses int."""
    if isinstance(value, bool):
        return None
    if isinstance(value, int) or isinstance(value, float):
        if value != value or value in (float("inf"), float("-inf")):
            return None
        if value < 0:
            return None
        return value
    return None


def check_rows(rows, rep):
    numbers = []
    for index, row in enumerate(rows, 1):
        if not isinstance(row, dict):
            rep.fail("row %d" % index, "not an object: %r" % (row,))
            continue
        keyword = row.get("keyword")
        label = keyword if isinstance(keyword, str) and keyword.strip() else "(unnamed)"
        if label == "(unnamed)":
            rep.fail("row %d" % index, "no keyword on the row")
        if "volume" not in row:
            rep.fail("row %d" % index, "%s: no volume field at all" % label)
            continue
        value = usable_volume(row["volume"])
        if value is None:
            rep.fail("row %d" % index,
                     "%s: volume is not a usable number: %r" % (label, row["volume"]))
            continue
        numbers.append(value)
        rep.ok("row %d" % index, "%s: %s" % (label, value))
    return numbers


def audit(payload, rep):
    if not isinstance(payload, dict):
        return None
    rep.note("export header")
    header = payload.get("export")
    if not isinstance(header, dict):
        rep.fail("export", "no export header object on the file")
        header = {}
    check_source(header.get("source"), rep)
    check_retrieved(header.get("retrieved"), rep)
    check_method(header.get("method"), rep)
    check_scope(header.get("scope"), rep)

    rows = payload.get("rows")
    rep.note("rows")
    if not isinstance(rows, list):
        rep.fail("rows", "no rows list on the file")
        return rep
    if not rows:
        rep.fail("rows", "the file parses and declares no rows at all")
        return rep

    numbers = check_rows(rows, rep)
    rep.note("whole export")
    if not numbers:
        rep.fail("coverage",
                 "%d rows present and not one usable volume among them" % len(rows))
        return rep
    if len(numbers) < len(rows):
        rep.fail("coverage", "%d of %d rows carry no usable volume"
                 % (len(rows) - len(numbers), len(rows)))
    else:
        rep.ok("coverage", "%d of %d rows carry a usable volume"
               % (len(numbers), len(rows)))
    if len(numbers) >= 3 and all(float(n) % 10 == 0 for n in numbers):
        rep.warn("granularity",
                 "every value is a multiple of 10 - check it is the tool's rounding")
    else:
        rep.ok("granularity", "values are not uniformly rounded")
    return rep


def load(path):
    """Return (payload, error_message)."""
    try:
        with open(path, encoding="utf-8") as handle:
            return json.load(handle), None
    except FileNotFoundError:
        return None, "no such file: %s" % path
    except IsADirectoryError:
        return None, "that is a directory, not an export: %s" % path
    except UnicodeDecodeError as exc:
        return None, "not UTF-8 text: %s" % exc
    except json.JSONDecodeError as exc:
        return None, "not valid JSON: %s" % exc
    except OSError as exc:
        return None, "could not be read: %s" % exc


def main(argv):
    if len(argv) < 2:
        print("usage: volume_provenance.py export.json | --selftest")
        return 4
    if argv[1] == "--selftest":
        raise RuntimeError("selftest: deliberate fault, the wrapper should catch this")

    payload, error = load(argv[1])
    if error is not None:
        print("INPUT UNUSABLE: %s" % error)
        return 4
    if not isinstance(payload, dict):
        print("INPUT UNUSABLE: top level is %s, expected one object with "
              "'export' and 'rows'" % type(payload).__name__)
        return 4

    rep = Report()
    audit(payload, rep)
    print("provenance audit of %s" % argv[1])
    rep.dump()
    if rep.fails:
        print("verdict: NOT AUDITABLE - %d failed, %d warned" % (rep.fails, rep.warns))
        return 3
    if rep.warns:
        print("verdict: AUDITABLE WITH WARNINGS - %d warned" % rep.warns)
        return 2
    print("verdict: AUDITABLE - every check passed")
    return 0


if __name__ == "__main__":
    # main() is wrapped so that no exception can ever surface as a finding.
    # SystemExit is re-raised untouched: it carries main()'s own return code.
    try:
        STATUS = main(sys.argv)
    except SystemExit:
        raise
    except BaseException as exc:
        print("INTERNAL ERROR: %s: %s" % (type(exc).__name__, exc))
        STATUS = 5
    sys.exit(STATUS)

Six runs follow, against five fixtures plus the deliberate fault. They land on five codes and never on 1; two land on 3, intentionally, because nonsense and nothing earn the same verdict. The fixtures use invented keywords and placeholder figures throughout. The transcripts were produced on 2026-10-02, and the retrieved check runs against the day you run it, so run one gains a staleness warning once that fixture date is 180 days old.

Run one: a header and rows that answer all four questions

{
  "export": {
    "source": "Own account export, Google Keyword Planner",
    "retrieved": "2026-09-20",
    "method": "twelve-month average as the tool labels it in this export, modelled not counted",
    "scope": "United Kingdom, English, Sep 2025 to Aug 2026"
  },
  "rows": [
    {"keyword": "keyword alpha", "volume": 1483},
    {"keyword": "keyword beta", "volume": 94},
    {"keyword": "keyword gamma", "volume": 2217}
  ]
}
$ python3 volume_provenance.py clean.json
provenance audit of clean.json
export header
  [ OK   ] source      Own account export, Google Keyword Planner
  [ OK   ] retrieved   2026-09-20
  [ OK   ] method      twelve-month average as the tool labels it in this export, modelled not counted
  [ OK   ] scope       United Kingdom, English, Sep 2025 to Aug 2026
rows
  [ OK   ] row 1       keyword alpha: 1483
  [ OK   ] row 2       keyword beta: 94
  [ OK   ] row 3       keyword gamma: 2217
whole export
  [ OK   ] coverage    3 of 3 rows carry a usable volume
  [ OK   ] granularity values are not uniformly rounded
verdict: AUDITABLE - every check passed
exit code: 0

Run two: the same file with every value rounded

Nothing failed, so the file is still auditable. Uniform rounding earns a look: it is equally the signature of a tool that rounds and of a person who did.

$ python3 volume_provenance.py rounded.json
provenance audit of rounded.json
export header
  [ OK   ] source      Own account export, Google Keyword Planner
  [ OK   ] retrieved   2026-09-20
  [ OK   ] method      twelve-month average as the tool labels it in this export, modelled not counted
  [ OK   ] scope       United Kingdom, English, Sep 2025 to Aug 2026
rows
  [ OK   ] row 1       keyword alpha: 1500
  [ OK   ] row 2       keyword beta: 90
  [ OK   ] row 3       keyword gamma: 2200
whole export
  [ OK   ] coverage    3 of 3 rows carry a usable volume
  [ WARN ] granularity every value is a multiple of 10 - check it is the tool's rounding
verdict: AUDITABLE WITH WARNINGS - 1 warned
exit code: 2

Run three: the export a model produced

{
  "export": {
    "source": "the model",
    "retrieved": "2026",
    "method": "",
    "scope": "UK"
  },
  "rows": [
    {"keyword": "keyword alpha", "volume": 1500},
    {"keyword": "keyword beta", "volume": "about 90"},
    {"keyword": "keyword gamma", "volume": true}
  ]
}
$ python3 volume_provenance.py generated.json
provenance audit of generated.json
export header
  [ FAIL ] source      names a generator, not a measurer: 'the model'
  [ FAIL ] retrieved   not an ISO date (YYYY-MM-DD): '2026'
  [ FAIL ] method      no collection method described: ''
  [ FAIL ] scope       no market, language or period given: 'UK'
rows
  [ OK   ] row 1       keyword alpha: 1500
  [ FAIL ] row 2       keyword beta: volume is not a usable number: 'about 90'
  [ FAIL ] row 3       keyword gamma: volume is not a usable number: True
whole export
  [ FAIL ] coverage    2 of 3 rows carry no usable volume
  [ OK   ] granularity values are not uniformly rounded
verdict: NOT AUDITABLE - 7 failed, 0 warned
exit code: 3

Row three catches people. In Python bool subclasses int, so a lazy isinstance(value, int) takes True for a search volume. Booleans are rejected first.

Run four: the file that parses and says nothing

This is the run that matters, and the one a careless checker gets wrong. The header is immaculate, the rows are correctly structured, not one carries a value. Read a missing field as an empty string, or as zero, and you ship an export in which nothing was measured.

{
  "export": {
    "source": "Own account export, Google Keyword Planner",
    "retrieved": "2026-09-20",
    "method": "twelve-month average as the tool labels it in this export, modelled not counted",
    "scope": "United Kingdom, English, Sep 2025 to Aug 2026"
  },
  "rows": [
    {"keyword": "keyword alpha", "volume": null},
    {"keyword": "keyword beta", "volume": null},
    {"keyword": "keyword gamma", "volume": null}
  ]
}
$ python3 volume_provenance.py empty.json
provenance audit of empty.json
export header
  [ OK   ] source      Own account export, Google Keyword Planner
  [ OK   ] retrieved   2026-09-20
  [ OK   ] method      twelve-month average as the tool labels it in this export, modelled not counted
  [ OK   ] scope       United Kingdom, English, Sep 2025 to Aug 2026
rows
  [ FAIL ] row 1       keyword alpha: volume is not a usable number: None
  [ FAIL ] row 2       keyword beta: volume is not a usable number: None
  [ FAIL ] row 3       keyword gamma: volume is not a usable number: None
whole export
  [ FAIL ] coverage    3 rows present and not one usable volume among them
verdict: NOT AUDITABLE - 4 failed, 0 warned
exit code: 3

Run five: a file that is not JSON at all

{"export": {"source": "Own account export",, "rows": []}
$ python3 volume_provenance.py broken.json
INPUT UNUSABLE: not valid JSON: Expecting property name enclosed in double quotes: line 1 column 44 (char 43)
exit code: 4

Run six: proving the crash path cannot pose as a finding

The --selftest flag raises an exception on purpose inside main(), so you can watch the wrapper catch it and land on a code no export reaches.

$ python3 volume_provenance.py --selftest
INTERNAL ERROR: RuntimeError: selftest: deliberate fault, the wrapper should catch this
exit code: 5

What it does not do matters as much. It makes no network calls, so it cannot say a number is wrong, only that it cannot be traced. The method check is weakest: twenty characters and no evasive phrase, so an empty but wordy description gets through.

Three demand signals that are not a volume

Three signals you can get without a subscription do carry provenance, and not one is a volume. They earn their place because a deliverable still has to put topics in an order, and ordering is the job they do. Each needs its limit written beside it in the file.

How crowded the first screen is. Page-one composition measures saturation rather than demand, and it is the cheapest of the three to read. Across the neighbouring task families recorded for doctor-seo.net on 20 September 2026, meaning keyword clustering, alt text, content briefs and reporting, page one was held almost entirely by vendor pages, with one programmatic domain taking four of nine results on clustering-method queries. That is a single dated look at pages that change and are personalised, so it is not re-runnable the way a stem is. It says the field is crowded, and nothing about how many people are looking.

Who is already doing the task. Adoption data answers a different question from demand data: it counts practitioners, not searches. In Keyword.com’s State of AI in SEO 2026, 1 January 2026, n=97 usable responses from self-selected respondents, 56% of them in teams of one to five, keyword research and clustering sits at 68%, in the top band with briefs, outlines and drafting, and one respondent of the 97 called any workflow fully automated. No margin of error is published, and with 97 self-selected responses the figures are an ordering at best: they are not shares of the profession, and small gaps between tasks are not differences. Read loosely, one thing still stands out: among the tasks respondents hand over most, keyword research is the one whose core number no model holds.

How deep the suggestion tree runs. Google offers at most ten completions for a stem, and the count coming back is informative by itself. A stem filling nine or ten slots on topic has people behind it; a stem managing two before it slides into an unrelated sense of the same words is padded out. This is ordinal evidence only: present or absent, never how many. The slides are the clearest part of the Google Autocomplete probes harvested for doctor-seo.net on 20 September 2026.

Stem probed, 20 September 2026 On-topic suggestions Where it drifted
ai for keyword research 10 of 10 no drift
gemini for seo 6 of 10 “gemini seoul”, agencies named Gemini
copilot for seo 2 of 5 aviation co-pilots
local llm for seo 1 of 5 the Master of Laws degree
gemini gem for seo 1 of 4 gemstones for the star sign

Run the method on the phrase at the top of this page and it is unflattering. On 20 September 2026 can chatgpt do keyword research produced three suggestions: itself, how to do keyword research and how long does keyword research take. It gets typed and carries no tree. The stem filling all ten slots is ai for keyword research, so the demand is for the task with the engine left blank.

No search volume for any of these terms was obtainable from a named source, so none is published here and none was invented. The fixture figures above are placeholders on invented keywords, there to exercise the script.

Connecting a source makes a number sourced, not correct

There is one repair for an invented volume and it is procedural rather than clever: put something that measures on the other end of the connection, and write down the day you pulled. The model is then no longer the origin of any figure, and the work it is genuinely good at happens on top of data that already answers the four questions.

Three connection shapes appear in the record. A custom GPT with Actions calls a keyword API directly; Paul Shapiro published the worked example at Search Wilderness on 28 November 2023, and building one is a lesson of its own. MCP connectors reach the paid suites, the Ahrefs connector listed as Anthropic-verified with 61 tools on 20 September 2026 and the Semrush MCP documentation last updated 5 August 2026; both return the provider’s own figure, which is where you go for its method and date. Third is demand already measured on your own property, covered in the Search Console and GA4 lesson.

Now the claim that has to stay small. A connected source converts an invented number into a sourced one, not into a correct one: a platform volume is a provider’s modelled estimate with a stated window, not a census of searches. ContextBolt published a week of live SEO work run through Claude on 23 June 2026 and recorded the model treating third-party estimates as factual data, then pushing a keyword forward after its own difficulty reading had already ruled it out. One practitioner, one site, seven days, author given only as “David”: a case, not a measurement.

That finding is not only about the model: promoting an estimate to a fact is something people do as readily as software, so the test runs after the connector is in place, not instead of it.

Underneath sits a real argument the sources cited here cannot settle. One camp holds that a platform’s modelled volume is the best measurement available and the only practical basis for prioritising work, carrying a documented method, a date and a provider who can be held to it. The other holds that a figure from a provider who does not publish the model is an estimate in the clothes of a count, and that the honest unit is relative demand rather than absolute monthly searches. No study in the sources cited here compares tool estimates against a disclosed ground truth, so neither camp can show its case. Both agree on what this lesson asks.

What a research mode or an agent actually buys you

A research mode or an agent buys you a trail, not a dataset. Text arrives with links attached, or steps run inside tools you already pay for; both are genuine gains in collection. Neither creates a measurement where nobody took one. On 20 September 2026 the stem chatgpt deep research seo returned an empty autocomplete tree, a fact about what people type rather than a verdict on the feature.

The trail is worth having, because it turns the four questions into something you can act on. Open a citation and two fields usually fill themselves in. More often the cited page quotes somebody quoting somebody, so you move a step along and nothing closes. Walk it to the end: either a name appears beside a measurement, or the trail runs out, and running out is the answer.

Common mistakes

  • Treating “it came from a tool” as the whole answer. The source field is one of four. A tool name with no export date, no stated window and no market is three answers short. The fix: fill all four, or the row is not finished.
  • Setting the bar at correct instead of at traceable. You cannot check a volume against anything, which is why the verdict is auditable or not. The fix: judge a row on whether somebody else could repeat the pull, never on whether it feels about right.
  • Trusting a checker you have never fed a broken file. A validator that has only seen good inputs is untested, and the expensive failure is the quiet pass. The fix: keep a deliberately empty export and a malformed one, and confirm each lands on the code you expect.
  • Averaging two providers into one column. Two modelled estimates built over different windows are not two readings of one quantity, so their mean belongs to nobody. The fix: one column per provider, named at the top, and compare orderings rather than values.
  • Letting the connector end the argument. A sourced figure is still somebody’s model, and the habit of promoting an estimate to a fact survives the wiring. The fix: run the four questions on what the connector returns too.

The short version

  • The verdict a figure gets is auditable or not auditable, never true or false. Could a second person repeat the pull from your four answers alone?
  • The four answers are who measured it, when it was pulled, what was counted and who it is about. Four, or the row does not ship.
  • No named source would give a volume for any of these terms, so none is published here; the fixture figures are placeholders on invented keywords.
  • No general-purpose chatbot holds a search-volume dataset, so a monthly figure it produces is a completion, not a reading. Si Quan Ong called model-quoted volumes “simply guesses” for Ahrefs on 30 October 2025, in a piece that publishes no sample size.
  • The test is indifferent to where a number came from: an unprovenanced chatbot figure and an unprovenanced tool export fail it the same way.
  • Build the check so a broken input cannot read as a clean one: absence has to be a failure, and the crash path needs a code nothing else uses.
  • A connector turns an invented figure into a sourced one, which is a smaller claim than correct: it is still a provider’s model with a window and a date.
  • With no named source available, ship the gap and name the substitute: page-one composition measures saturation, adoption surveys count practitioners, autocomplete depth is ordinal.

Frequently asked questions

Does this apply to keyword difficulty and cost-per-click too?

Yes, and the test does not change. Difficulty and cost-per-click are modelled outputs of the same products, so they need the same four answers, and a model with nothing connected composes them as it composes a volume. ContextBolt’s trial of 23 June 2026 is the documented case, one site and one week: a difficulty score was on screen and the recommendation went against it.

Can ChatGPT do keyword research?

Everything except the part that needs a dataset. It is quick and checkable at generating variants, grouping a list you paste in, labelling intent and turning the result into a brief. Volumes, difficulty and competition are not inside it, so it produces those by completion. Point it at a keyword API through a custom GPT with Actions and the figures become the API’s, with its date and window attached.

The script says my export is not auditable, but I know the numbers are fine. Now what?

Fill in what is missing rather than overriding the verdict. Not auditable is about the paperwork, not the figures: nobody downstream can repeat your pull. Go back to the tool, record the export date and the window it documents, write market and language into the header, run it again. If that information does not exist, that is the real problem.

Sources

  • Google Autocomplete probes harvested for doctor-seo.net, 20 September 2026, via suggestqueries.google.com. Ordinal evidence only.
  • Result-page composition recorded for doctor-seo.net on 20 September 2026 across the keyword clustering, alt text, content brief and reporting families: one programmatic domain held four of nine results on clustering-method queries. One observation, no sample size.
  • Ahrefs connector for Claude, Anthropic-verified with 61 tools, checked 20 September 2026. Connector inventories move in weeks, so the check date is part of the fact.
  • Semrush MCP documentation, last updated 5 August 2026.
  • ContextBolt, author given only as “David”, “Claude SEO Experiment: A Week of Running My Real SEO”, 23 June 2026. One practitioner, one site, seven days; no control, no sample size.
  • Keyword.com, “State of AI in SEO 2026”, 1 January 2026. n=97 usable responses, self-selected, 56% in teams of one to five; keyword research and clustering 68%; fully automated 1%. No margin of error.
  • Paul Shapiro, “Using ChatGPT’s Custom ‘GPTs’ for SEO and Keyword Research”, Search Wilderness, 28 November 2023, updated 27 December 2023.
  • Si Quan Ong, “AI can’t replace SEO tools, but it can use them”, Ahrefs, 30 October 2025. Model-quoted volumes called “simply guesses”. An argument piece, not a study; no sample size published. Named and dated here, not linked.

Continue the course

The previous lesson is the routing table for which engine suits which task. The index for this level is AI for SEO: Using the Models as Tools, and every level sits on the course hub.