SEO TOOLS
XML Sitemap Best Practices: What Google Reads and What It Throws Away
7 min read · ToolsBay editorial · Published · Updated
Just want to do it now?
Turn a list of URLs into a valid sitemap.xml for Google and Bing.
A sitemap does one job: it hands a crawler a list of URLs it might not otherwise find. That is the entire mechanism. It does not make pages rank, it does not force anything into the index, and Google's own documentation is blunt about the ceiling: "Submitting a sitemap is merely a hint: it doesn't guarantee that Google will download the sitemap or use the sitemap for crawling URLs on the site."
Most advice treats the file as a lever. It is closer to a table of contents left on the counter — useful when the building is large or the hallways are confusing, pointless when there are three rooms and every door is open.
What follows is what is actually true about the format, most of which contradicts the boilerplate.
The two limits, and the one that never bites
Every guide quotes the same pair: 50,000 URLs and 50 MB uncompressed per file. Both are real, both come from the sitemaps.org protocol, and Google restates them. What nobody mentions is that the byte limit is almost unreachable.
Here is one entry with every optional field filled in:
<url>
<loc>https://example.com/products/running-shoes-mens</loc>
<lastmod>2026-09-10</lastmod>
<changefreq>weekly</changefreq>
<priority>0.8</priority>
</url>With a 60-character URL, that entry is 192 bytes. Fifty thousand of them come to about 9.2 MB — under a fifth of the cap. Strip the three optional fields and the same 50,000 URLs weigh about 4.4 MB. Push the URLs out to 120 characters and keep all four fields and you land near 12 MB.
To hit 50 MB inside 50,000 entries, each entry would have to average 1,000 bytes, which means URLs roughly 900 characters long. If your URLs look like that, you have a different problem.
So the practical rule is simpler than the one everyone repeats: split on the count. The byte limit is a footnote, and the only time it is worth a thought is with gzipped sitemaps, where the size that counts is the uncompressed one.
changefreq and priority are ignored
Not "weighted lightly". Ignored. Google's sitemap documentation says it in five words: "Google ignores <priority> and <changefreq> values."
They survive because they were in the 2005 protocol and most generators written since have emitted them by default. Keeping both costs about 99 bytes per URL — roughly 4.7 MB across a full 50,000-URL file — spent on two elements nothing reads. There is no penalty for including them. There is also no reason to.
The sitemap generator on this site defaults changefreq to weekly and leaves priority blank, and both can be set to nothing at all. A sitemap containing only <loc> elements is complete and valid.
lastmod is the one field that carries weight, and the one most sites ruin
Google's phrasing: "Google uses the <lastmod> value if it's consistently and verifiably accurate." Both qualifiers do work.
Verifiable is the easy half. Google has the page. If your sitemap claims a URL changed yesterday and the content it fetches is identical to what it crawled in March, the claim was checkable and it failed.
Consistent is the half that costs sites the field. The judgement is made per site, not per URL. A deploy pipeline that stamps today's date on every URL on every build — a very common default, and the behaviour of more than one popular sitemap plugin — teaches Google that on this domain lastmod means "we deployed", not "this page changed". Once the field is discounted for a site, an honest date on the one page that genuinely changed does not buy much back, and not quickly.
That is why a wrong lastmod is worse than no lastmod. Omitting it costs you nothing. Faking it costs you the field. If your build cannot tell which pages actually changed, leave the element out and lose no ground.
Where the file sits decides which URLs it may contain
From the protocol: "The location of a Sitemap file determines the set of URLs that can be included in that Sitemap." A sitemap at /shop/sitemap.xml covers URLs under /shop/. One at the domain root covers the site. And every URL listed must use the same protocol and the same host as the sitemap itself.
"Same host" is string matching, not what a person means by "the same site". https://example.com and https://www.example.com are different hosts. So are the http:// and https:// forms. A sitemap served from the www hostname that lists bare-domain URLs is a quiet, common failure — the file parses, the entries look right, and they are out of scope.
There are two ways to widen that scope. Google accepts URLs across properties you have verified when the sitemap is submitted through Search Console. The protocol separately treats a Sitemap: reference in robots.txt as evidence you control the host serving that robots.txt. That line is also how a crawler that has never heard of your site finds the list in the first place; the robots.txt generator writes it into the file for you.
Index files, and the nesting that silently fails
Past 50,000 URLs you split into several sitemaps and list those in a sitemap index — a different root element, up to 50,000 child entries, and its own 50 MB cap:
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://example.com/sitemap-products-1.xml</loc>
<lastmod>2026-09-10</lastmod>
</sitemap>
</sitemapindex>Two levels — one index, 50,000 sitemaps, 50,000 URLs each — reaches 2.5 billion URLs, so the structure never needs to go deeper. Which is fortunate, because it cannot: a sitemap index may list sitemaps only, never another sitemap index. Google's documentation states this directly, and Search Console reports a nested index as an error rather than following it.
The reason it catches people is that nesting is the obvious move. A large site shards by section, each section owns its own index, and someone wraps those indexes in a parent index at the root. That parent is the one file Google reads, and it contains nothing Google will follow.
Escaping is the usual reason a sitemap is rejected
XML has five characters that cannot appear raw in text: &, <, >, " and '. In a sitemap, the one you actually meet is the ampersand, because it separates query parameters.
<loc>https://example.com/search?q=shoes&sort=price</loc>That document is not well-formed. A parser stops at the &, and because XML is not error-tolerant, one bad URL invalidates the whole file rather than its own entry. The fix is the entity:
<loc>https://example.com/search?q=shoes&sort=price</loc>Separately, and beforehand, anything outside ASCII has to be percent-encoded per RFC 3986: /café becomes /caf%C3%A9. Two encodings, applied in that order, and hand-assembled sitemaps routinely miss one of them.
When a sitemap helps, and when it changes nothing
It helps on large sites, where crawling everything by following links takes real time. On new sites with few inbound links, where a crawler has almost nothing to follow. On pages reachable only through faceted navigation or an internal search box. On a section linked from exactly one place, deep in the tree.
It changes nothing on a 30-page site with an ordinary navigation that is already indexed. The crawler found every page on day one by following links. A list of the same URLs adds no information it does not already have.
And the thing people most often expect a sitemap to fix, it does not: "Discovered – currently not indexed" in Search Console. That status means Google already knows the URL exists and has chosen not to crawl or index it yet. Listing it again asserts something Google has already accepted. Discovery was never the problem there.
One practical note that outdated guides still get wrong: the ping endpoint is gone. Google announced its removal in June 2023, and requests to it now return 404. Bing retired its equivalent as well. Any plugin still firing that request on every publish is doing nothing at all. Submission is Search Console, or the robots.txt line, or neither.
The short version
- List absolute, canonical URLs that return 200 and are indexable. A URL that redirects, 404s, or carries a
noindextag sends two contradictory signals — check the tags with the meta tag generator. - Drop
changefreqandpriority. - Include
lastmodonly if it is genuinely per page. Otherwise omit it. - Escape the ampersands, then look for anything non-ASCII.
- Split at 50,000 URLs. Ignore the 50 MB figure.
- Put the file at the root, reference it from robots.txt, submit it once.
Everything after that is the crawler's decision, and no element in the file changes it.