Skip to content
Doctor SEO

XML Sitemap Checker: Validate Your Sitemap and Test Every Crawler

A free XML sitemap checker that validates your sitemap against Google's documented limits, opens every sub-sitemap of an index and shows whether Googlebot, Bingbot and 12 AI crawlers can fetch it.

Paste a sitemap URL, or just your domain, and this XML sitemap checker reads the file the way Google’s documentation says a sitemap must be built: 50MB and 50,000 URLs at most, UTF-8, the exact sitemap namespace, absolute URLs, valid dates. Then it does two things most validators skip. It opens the sub-sitemaps of a sitemap index, and it fetches your sitemap again as 14 different crawlers, from Googlebot and Bingbot to GPTBot, ClaudeBot and PerplexityBot, so you can see who actually gets in.

Paste the sitemap URL, or just the domain and the checker will find it through robots.txt.

Requests are sent with this crawler’s user agent and robots.txt is read with its rules. Every crawler is also compared in the access table.

The most useful finding is often not an XML error at all. When we ran the checker on doctor-seo.net itself on 10 October 2026, every sampled page answered a browser and the ClaudeBot user agent with 200, while the GPTBot user agent got 429 Too Many Requests on all 10 sampled pages, then 200 again a few seconds later. The sitemap was valid. A rate limit in front of the site was treating one crawler differently from the others, and no syntax validator would have shown it.

What the XML sitemap checker tests

The doctor-seo.net XML sitemap checker runs about 120 checks and names each failure the way the Sitemaps report in Google Search Console names it, so a finding here maps onto an error you may already have seen there: “Couldn’t fetch”, “Path mismatch”, “Nested sitemap indexes”, “Invalid date”. The checks fall into seven groups.

Group What it checks Example of a blocking error
Fetch HTTP status, redirects, response time, Content-Type, gzip and decompression, empty file, an HTML page served instead of XML The sitemap URL returns 404, 403 or 5xx
Format UTF-8 (declared and actual bytes), whitespace before the XML declaration, DOCTYPE, well-formed XML with line and column, exact namespace, undeclared prefixes, misspelled extension namespaces, repeated tags A parse error, after which Google cannot read the rest of the file
URLs Missing or empty <loc>, relative URLs, URLs of 2,048 characters or more, unencoded characters, # fragments, tracking parameters, www and http/https mismatch, URLs outside the sitemap’s folder, duplicates URLs on another host (“URL not allowed”)
lastmod Coverage, W3C Datetime format, real calendar dates, future dates, one date on every URL, dates that look like the time of the request A date in the wrong format (“Invalid date”)
Sitemap index Incomplete child URLs, children on another host or outside the index’s folder, nested indexes, children that fail to load or redirect An index that lists another index
Extensions Image, video, news and hreflang entries: required tags, limits, deprecated tags, hreflang codes and return links A video without a thumbnail URL
Crawling robots.txt for the selected crawler (sitemap, sampled URLs, sub-sitemaps), a live test of up to 10 URLs from the sitemap (status, redirect, noindex, canonical) and the crawler access table A Disallow rule that matches the sitemap URL

You can also paste raw XML, up to 5MB, to check a sitemap that is not public yet. Add the URL it will live at and the host and folder checks run as well.

How to read the results

The checker sorts findings into four levels, and only the first one is a reason for Google Search Console to reject a sitemap or skip its URLs. Errors are what Search Console reports as errors or what stops Google using the URLs: a sitemap that returns 404, broken XML, a wrong namespace, an invalid date. Warnings leave the sitemap valid but cost you something, such as URLs that redirect, a sitemap missing from robots.txt or URLs on another host. Improvements are optional and worth doing, with lastmod coverage at the top. Good to know is information, such as a gzip file or the number of image entries.

Each finding in the report shows a count, up to five examples from your file and a link to the Google documentation it rests on. “Copy report” copies the whole result as plain text for a ticket, and the share link reruns the same check with the same crawler selected.

A green “Valid sitemap” verdict means nothing blocking and no warnings. It does not mean Google will crawl or index the URLs. Google’s own guide is explicit that “submitting a sitemap is merely a hint: it doesn’t guarantee that Google will download the sitemap or use the sitemap for crawling URLs on the site.” What happens after discovery is covered in our post on the crawl and index timings Google has put on record.

Testing whether Googlebot, Bingbot and AI crawlers can fetch your sitemap

The “Crawl as” menu sends every request with the chosen crawler’s user agent and reads robots.txt with that crawler’s rules. The matching follows RFC 9309, the IETF standard for the Robots Exclusion Protocol: a group that names the crawler wins over the * group, and inside a group the longest matching path wins. Google’s robots.txt documentation describes the same rule: “crawlers use the most specific rule based on the length of the rule path.”

Under the main report, the crawler access table repeats the test for all 14 crawlers at once and compares each one with a normal browser. For every crawler it shows whether robots.txt allows the sitemap, the sampled URLs and the sub-sitemaps, and what the server answered to that user agent.

Group Crawlers in the table What the operator uses them for
Search engines Googlebot, Bingbot, Applebot Web search results (Bing’s index also feeds Copilot)
AI search OAI-SearchBot, Claude-SearchBot, PerplexityBot The index behind answers in ChatGPT search, Claude and Perplexity
AI assistants, on user request ChatGPT-User, Claude-User, Perplexity-User Fetching a page a user asked the assistant to read
AI training GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent Collecting data for model training

Google-Extended and Applebot-Extended are robots.txt tokens, not separate crawlers: Google and Apple honour the rule but fetch with Googlebot and Applebot. The table marks them that way instead of inventing a user agent.

The checker does not tell you which AI crawlers to allow. Whether to block AI training or AI answering bots is a business decision the industry is split on, and doctor-seo.net has not taken a position on it. What the table shows is whether what you block is what you meant to block. That is where most surprises come from: a CDN switch such as “block AI bots” refusing a user agent at the server while robots.txt says Allow, or a Disallow rule written years ago that now catches a sub-sitemap.

The caveat: an imitation is not the real crawler

Every test request comes from the doctor-seo.net server, not from the published IP ranges of Google, Microsoft, OpenAI or Anthropic. A firewall that verifies crawlers properly will block the imitation and still let the real crawler in. Google documents exactly that kind of check in Verify requests from Google crawlers and fetchers: a reverse DNS lookup, or matching the IP address “against the list of published Google IP addresses.”

So a refusal that hits only the Googlebot user agent on a large site is often a firewall doing its job. Confirm it with the URL Inspection tool in Search Console, which fetches as the real Googlebot. A refusal that hits every AI crawler and no search crawler is more often a CDN rule. A 429, like the one GPTBot’s user agent got on this site, is a rate limit; if you slow crawlers on purpose, our post on Google’s Retry-After documentation covers what Google now says about doing it.

Google’s sitemap limits at a glance

Google’s sitemap limits fit in one table, and every row below links to the page that states it. The checker tests each of them.

Limit Value Source
Size of one sitemap 50MB uncompressed Build and submit a sitemap
URLs in one sitemap 50,000 Build and submit a sitemap
Encoding UTF-8 Build and submit a sitemap
URLs Fully qualified, absolute Build and submit a sitemap
Scope “A sitemap affects only descendants of the parent directory” Build and submit a sitemap
Sitemaps in one index 50,000 <loc> tags Manage large sitemaps
Index files per site in Search Console 500 Manage large sitemaps
Where child sitemaps live Same directory as the index or lower Manage large sitemaps
Length of a <loc> URL Fewer than 2,048 characters sitemaps.org protocol
Images per <url> 1,000 Image sitemaps
News sitemap Articles from the last two days, up to 1,000 news:news tags News sitemaps

lastmod: the one optional tag Google reads

Google reads <lastmod> and ignores <priority> and <changefreq>. Its sitemap guide, last updated on 8 July 2026, says it in one line: “Google ignores <priority> and <changefreq> values.” The same guide says Google uses lastmod only when the value is consistently and verifiably accurate, for example when it matches the page’s actual last modification.

Google spelled out what counts as a change when it announced the end of the sitemaps ping endpoint on 26 June 2023. A CMS changing “an insignificant piece of text in the sidebar or footer” needs no new date, but “if you changed the primary text, added or changed structured data, or updated some links, do update the lastmod value.”

That is why the checker looks past “present or missing”. It flags four patterns that teach Google to distrust the dates:

  • Every URL has the same lastmod. The generator is almost certainly writing the moment the file was built.
  • 90% or more of the dates fall in the 15 minutes before the check. The sitemap prints “now” instead of each page’s last real change.
  • Dates in the future. Usually a time-zone error or scheduled content.
  • An index lastmod older than the newest URL in that child sitemap. The index is not updated when the child changes.

A missing lastmod is reported as an improvement, not an error. Leave it out on pages whose real change date you cannot know, such as a home page or a category page that only lists other pages. Everywhere else, use W3C Datetime: 2026-10-10, or 2026-10-10T09:30:00+02:00 with a time zone.

Sub-sitemaps: why the checker opens them

A sitemap index is only as good as the sitemaps it lists, and a valid index pointing at a broken child still loses those URLs. When you check an index, the checker fully analyses the first 10 child sitemaps, then requests the rest, up to 200, in batches of 25, and confirms that each one answers 200 and really is a sitemap. Every row has a “Check” button that runs the full analysis on that child alone.

The failures this finds are ordinary: a child sitemap removed by a plugin update that now redirects to the home page, a child caught by a robots.txt Disallow written for something else, a child on the www host while the index is not. The checker also compares URLs across the children it analysed, because the same URL in a post sitemap and a page sitemap is a generator bug. In a test in October 2026 against a national newspaper’s index of 2,101 child sitemaps, the checker reached the 200-child cap in about 11 seconds.

Common sitemap mistakes and how to fix them

  1. The robots.txt Sitemap: line points at an old location. Fix the line. Google’s robots.txt documentation says the field “isn’t tied to any specific user agent” and “must be a fully qualified URL, including the protocol and host.”
  2. An unescaped & in a URL. It breaks the XML at that point. The sitemaps.org protocol says “all data values in a Sitemap must be entity-escaped”, so write &amp;.
  3. URLs that redirect, carry noindex or canonicalise elsewhere. List only the final, canonical URL. The checker’s live sample of up to 10 URLs catches all three.
  4. A blank line before <?xml. Search Console reports it as “Leading whitespace”. The usual cause is a blank line before <?php in a theme or plugin file.
  5. lastmod set to the time the sitemap was generated. Output each page’s real modification date, or leave the tag out.

What this checker cannot tell you

The XML sitemap checker tests the file and what a crawler gets when it asks for it. It cannot see inside Google.

  • Whether Google has read your sitemap, and when. Only the Sitemaps report in Search Console shows Google’s last read date and its count of discovered URLs.
  • Whether the URLs will be indexed. A sitemap helps discovery; indexing depends on the pages.
  • What the real crawlers see from their own IP addresses. See the caveat above.
  • Every URL in a big sitemap. The live test samples up to 10 URLs spread across the file, and child sitemaps over 10MB are skipped in the combined run (each can still be checked on its own).
  • Feeds in detail. RSS 2.0 and Atom 1.0 feeds are recognised as valid sitemap formats but not analysed tag by tag.

The short version

  • Google limits one sitemap to 50MB uncompressed or 50,000 URLs, and one index to 50,000 sitemaps.
  • Google ignores priority and changefreq and uses lastmod only when the dates are consistently accurate.
  • Update lastmod when the main text, the structured data or the links change, not when a footer does.
  • A sitemap index can be valid while one of its child sitemaps is broken; check the children too.
  • robots.txt rules and CDN rules are separate; a crawler can be allowed in one and refused by the other.
  • A test from an ordinary server imitates a crawler’s user agent, not its IP address, so confirm Googlebot problems with URL Inspection in Search Console.

FAQ

Is the XML sitemap checker free?

Yes. There is no account and no paywall. The only limits are a cap on checks per connection to keep the server available, and a 15-minute cache, so rechecking the same sitemap with the same crawler inside that window returns the stored result.

Why does my sitemap pass here but fail in Search Console?

Usually because the two are not looking at the same moment or from the same place. Search Console shows Google’s last read, which can be days old. A “Couldn’t fetch” in Search Console with a clean result here often means a firewall or CDN refuses Google’s real IP addresses, which this checker cannot test from its own server.

Do I still need to ping Google when my sitemap changes?

No. Google announced on 26 June 2023 that it was retiring the sitemaps ping endpoint. Submit the sitemap once in Search Console, list it in robots.txt and keep lastmod accurate.

Should I remove priority and changefreq?

They do no harm, and Google’s guide says it ignores both. Remove them if they make the file harder to maintain; keep them if another consumer of your sitemap reads them. Do not spend time tuning them for Google.

Can I check a sitemap index with hundreds of child sitemaps?

Yes. The checker analyses the first 10 in full and checks the rest, up to 200, for status and format. Use the “Check” button on any row to analyse that child sitemap completely.

Where to go next

This tool is part of the Doctor SEO tools section. The method behind it, how crawling, robots.txt and indexing fit together, belongs to the technical SEO level of the free SEO course. If a term in the report is new to you, the SEO glossary defines it.

Sources