AI for SEO: Using the Models as Tools
Using Claude, ChatGPT, Gemini and Perplexity to do SEO work — and above all, how to check what they produce. This is not how AI search cites you: that is Level 4.
By the end of this level you will be able to put Claude, ChatGPT, Gemini and Perplexity to work on real SEO jobs: keyword clustering, redirect maps, internal linking, log analysis, bulk metadata, reporting. More to the point, you will be able to prove the output is right before it reaches a live site. Generating the work is the cheap half. Checking it is the level.
Level 8 is about the models as tools. It is not about how AI search engines decide what to cite: that question belongs to Level 4 on GEO and AIO, and the two get conflated in almost every article written on the subject. “How do I get ranked inside ChatGPT” and “how do I do keyword research with ChatGPT” are different jobs resting on different evidence. If you came for the first one, Level 4 is the page you want; whether it is a separate discipline or ordinary SEO renamed is disputed in the industry, and Level 4 sets out that disagreement rather than settling it.
The most useful number behind this level is a small one. In Keyword.com’s State of AI and Automation in SEO (published 1 January 2026, n=97, self-selected, 56% of respondents in teams of one to five people), 1% of respondents had fully automated any SEO workflow. Ninety-nine per cent keep a person in the loop, and the reason given is quality: 57% said the output was not good enough, 36% said they did not trust its accuracy. Read those figures as an ordering, not a population estimate; n=97 is small and self-selected, and it is the only survey this level rests on.
One warning before you start. Everything here has a short shelf life. Connector inventories, context windows, tool versions and per-request pricing all moved during the month the research was done, so every lesson carries its date, and anything older than a quarter is a lead to re-check, not a fact to quote.
What this level covers
- Connect a model to real data: Search Console and GA4 over MCP, the paid suite connectors, a crawler, your CMS. Then work out which of those repay the setup time.
- Recognise the outputs no general-purpose model can be trusted on unaided: search volumes, difficulty scores, anything that looks like measurement and is generation.
- Build the checks that decide whether a batch of AI work ships: gold sets, similarity thresholds, acceptance sampling, a scored test set for intent labelling.
- Price a bulk run before committing to it, including the metered search charges that sit on top of token costs.
- Decide which tasks to delegate and which to keep, scored on the cost of a wrong output and on how cheaply it can be verified.
The inversion: generating is saturated, checking is empty
The web holds an enormous supply of “how to generate X with AI” and almost nothing on “how to check what it produced”. That asymmetry is why every task lesson in this level is a verification lesson rather than a generation lesson. Keyword.com’s survey (1 January 2026, n=97) shows where the volume already is: 77% of respondents use AI for briefs and outlines, 68% for keyword research and clustering, 66% for drafting. No published accuracy metric for any of those tasks could be found during this research.
Verification is also the skill that survives a model upgrade. A prompt ages out with the next release; a gold set of 100 hand-labelled items, a defect taxonomy and a stated release bar carry over. Gus Pelogia, quoted in Athens SEO’s practitioner round-up of 23 April 2026 (updated 3 June 2026, 13 named practitioners), names roughly 80% accuracy as the bar a prototype must clear before he will use it. Published bars of that kind are rare.
One check that does not work deserves naming early, because teams keep buying it. A study in the International Journal for Educational Integrity (Hadra, Cambridge and Mesbah, 2 February 2026) measured two commercial AI-content detectors at 0.69 and 0.61 overall accuracy, with near-zero recall on hybrid human-and-AI text. No sample size is stated alongside those accuracy figures, and that omission is worth knowing. Hybrid text is exactly what an SEO team produces, so detectors are least reliable on the only case that matters. Lesson 22 does the arithmetic on how many of your own writers a detector at that accuracy would wrongly accuse.
Every engine invents search volumes
No general-purpose chatbot has search volume data, and every one of them will hand you a number anyway. Si Quan Ong wrote for Ahrefs on 30 October 2025 that model-quoted volumes are “simply guesses”, quoting a practitioner comparing them to estimating that there are 47 million penguins living in your fridge. ContextBolt’s week-long trial of Claude on SEO work (23 June 2026, updated 11 September 2026, one unreplicated trial by one team) recorded the sharper failure: the model treated third-party estimates as factual data and kept pushing a keyword whose own difficulty score said it was unwinnable.
The fix is identical on Claude, ChatGPT, Gemini and Perplexity, which is why this level has one lesson on invented metrics rather than one per engine. Connect a real data source, or treat the number as fiction. That sameness is also why the level is organised by capability and task rather than by brand: keyword research with ChatGPT and keyword research with Gemini are roughly 70% the same argument with a different logo on top.
What a bulk AI run actually bills
Metered search charges are the cost nobody models, because they sit on top of token costs rather than inside them. The figures below come from the vendors’ own pricing and developer documentation, checked on 20 September 2026. They are list prices, not measurements, and they move, so re-check them before quoting a client.
| Surface | Free allowance | Metered charge after it | Charged on top of |
|---|---|---|---|
| Gemini, Grounding with Google Search (Gemini 3.x) | 5,000 requests per month | $14 per 1,000 requests | Tokens. Google states one submitted request may produce several billed search queries. |
| OpenAI web search tool | None documented | $10 per 1,000 calls | Tokens |
| Perplexity Sonar | None documented | $5–12 per 1,000 requests, by context size | $1 in / $1 out per million tokens |
| Perplexity Sonar Pro | None documented | $6–14 per 1,000 requests | $3 in / $15 out per million tokens |
| Batch processing (OpenAI, Google) | — | 50% discount on the standard rate | Not available for interactive work |
| DeepSeek off-peak | — | 50% discount | deepseek-flash at $0.15–0.30 in / $0.60–1.20 out per million tokens |
Run one worked example before building anything: URLs multiplied by the grounded searches each needs, put against the metered column. A 5,000-URL job issuing two searches per URL is 10,000 billed searches, which on several rows above costs more than the tokens do. Connectors change the arithmetic again: the Ahrefs connector alone exposed 61 tools when checked on 20 September 2026, and tool definitions consume context before you ask anything.
Which engine, and why the comparison pages fail
The comparison content the SEO industry assumes is in demand does not appear to exist. Three Google Autocomplete probes run for this course on 20 September 2026 (chatgpt vs gemini for, is gemini better than chatgpt for, perplexity vs chatgpt for) returned 30 suggestions between them, and SEO appeared zero times. The suggestions were coding, study, research, interior design and relationship advice. Probed separately, is claude better than chatgpt for returned ten suggestions with no SEO either.
What people do ask is engine-agnostic. The stems best ai for seo and ai for keyword research both sustained a full ten-suggestion tree on the same date. Autocomplete is ordinal evidence, not volume: a suggestion appears or it does not, and no figure comes with it. The shape is still clear enough to organise a level around, so Level 8 answers the agnostic question with one routing table rather than a matrix of pairwise comparisons. Grok, DeepSeek, Microsoft Copilot and locally run open models get rows in that table and no page of their own: their SEO stems collapsed within two or three suggestions into unrelated meanings.
What this level is not
Level 8 is not Level 4. Nothing here teaches you to be cited by an AI answer engine; that is Level 4. The one place the two genuinely touch is lesson 9, on how Claude reaches pages through Brave Search rather than Google.
Level 8 is also not a ruling on whether AI-written content is allowed. Google’s spam policies, last updated 28 August 2026, define scaled content abuse as many pages generated for the primary purpose of manipulating rankings and not helping users, “no matter how it’s created”. Production method is explicitly not the test, and Ahrefs found 9% of top-ranking pages to be at least 80% AI across 331,000 top-ten pages, published 27 July 2026. Where the line sits is argued in Level 6 and Level 7, not here.
And Level 8 is not a prompt listicle. Search Engine Land published a worked failure on 28 August 2026, reported by Will Scott: an assistant asked to build two new landing pages cloned the site’s homepage into both and changed only the title tags. Both pages recorded zero impressions and zero clicks, the homepage stayed at position nine or worse, and the same duplication recurred on a second site. Two sites and one author is not a sample, and saying so is part of using the case honestly. It is still a documented outcome, which is more than most writing here offers.
The 30 lessons in this level
Claude for SEO
- What Claude Can and Cannot Do for SEO
- Claude Skills for SEO: Install, Use, and Write Your Own
- Claude Code for Technical SEO
- Connecting Claude to Your Own Search Data: GSC and GA4 via MCP
- The Paid Suite Connectors: Ahrefs, Semrush, and Whether They Earn Their Setup
- Claude and Your Crawler: Screaming Frog’s MCP Server
- Claude and Your CMS: WordPress, Rank Math, and Write Access
- Claude’s Three Crawlers and Your robots.txt
- Getting Cited by Claude: The Brave Dependency
- Being Findable by Brave’s Web Discovery Project
- Claude Cowork for SEO
The other engines
- Which AI for Which SEO Task: A Routing Table
- Why Every Chatbot Invents Search Volumes, and What to Do Instead
- ChatGPT SEO Prompt Patterns, by Task
- Custom GPTs with Actions: Wiring Real Keyword Data into ChatGPT
- Gemini for SEO: Sheets, Workspace, and Where It Refuses
- Gemini’s Grounding with Google Search API for SEO Research
- Perplexity for SEO Research and Citation Checking
- Perplexity Pages and the Parasite SEO Question
Verification and delegation
- How to Validate an AI Redirect Map Before You Push It
- Choosing a Similarity Threshold for Automated Internal Linking
- Why AI-Content Detectors Should Not Be in Your Editorial Workflow
- A Go/No-Go Test for AI-Generated Pages at Scale
- What Breaks When You Drive a Crawler From an Agent
- Measuring Whether Your Keyword Clusters Are Any Good
- An Eval Harness for Intent Classification at Scale
- Acceptance Sampling for AI-Drafted Content
- The Delegation Triage: Which SEO Tasks to Give AI, and Which to Keep
- What Bulk AI SEO Work Actually Costs
- Anomaly Detection on Search Console Data
How to study this level
Take the three groups in order, but read lesson 28, the delegation triage, before you connect anything. It scores a task on two axes, what a wrong output costs and how cheaply it can be checked, and it will tell you several of the jobs you were about to automate should not be. Cheaper to discover on paper than after 400 pages are live.
Apply one gate before running anything at scale. Write down in advance what counts as a serious defect: a fabricated statistic, a fabricated citation, a wrong entity, an unverifiable claim stated as fact. Hand-label a sample yourself, compare the model’s labels against yours, and set the accuracy you will accept before you see the result. A bar chosen afterwards is a bar chosen to pass.
Run the failure modes deliberately. Agent-driven crawls break in documented ways: testing published by Rich Voller on 26 May 2026, updated 12 June 2026, recorded a crawl running 500 URLs in 20 seconds against an interface setting of one URL per second, and a progress indicator reaching 100% before the crawl had finished. That is one case from one agency with no sample size attached, so treat it as something to check, not as a rate. Reproducing it once on a staging site teaches more than reading about it.
Have the foundations in place first. Level 2 on technical SEO covers the crawling and indexing behaviour these workflows assume, and Level 5 on analytics and measurement covers the Search Console data lessons 4 and 30 read. An automated internal-linking pass on a site whose canonical tags are wrong just builds the wrong links faster.
Level 8 sits near the end of a longer path that starts with fundamentals and technical work and ends in measurement and strategy. If you arrived here directly, the full index with every level in its recommended order is on the free SEO course at doctor-seo.net. Come back with a live site and a real task in front of you: every lesson is written to be executed, not skimmed.