SEO TOOLS
Robots.txt Disallow Rules: What They Actually Control
8 min read · ToolsBay editorial · Published · Updated
Just want to do it now?
Write crawler allow/disallow rules without syntax mistakes.
robots.txt does not tell Google what to index. It tells crawlers what they may fetch. Those are two different things, and nearly every expensive mistake made with this file comes from treating them as one.
A URL you block with Disallow can still appear in search results. If another site links to it, Google knows the URL exists without ever fetching it, and it can list that URL on the strength of the link alone. You have seen the result: a bare URL or a title lifted from the anchor text, no description, and the line No information is available for this page. The page is indexed. There is just nothing behind it, because the one thing you did was stop Google looking.
Blocking is not deindexing
To take a page out of Google, you have to let Google crawl it and then tell it to go away — a <meta name="robots" content="noindex"> in the <head>, or an X-Robots-Tag: noindex response header for things that have no head, like a PDF or a JSON export. Google fetches the page, reads the instruction, and drops the URL from the index. The meta tag generator writes that tag alongside the rest of a page's head block.
This produces the trap that catches most people. They add noindex to a page, then add Disallow for the same path to be thorough. The Disallow wins, because it is evaluated first: Googlebot never requests the page, so it never sees the noindex, so the URL stays in the index as an empty stub indefinitely. The two directives do not stack. Pick one, and if you want the page gone, it is the noindex.
If the page is already indexed and you need it out today, the Removals tool in Search Console hides a URL from Google's results for roughly six months. It is a stopgap, not a fix — it buys time to get a real noindex in place.
Only one group applies to any given crawler
The protocol was a convention from 1994 until the IETF published RFC 9309 in 2022 and finally wrote it down. It is an IETF standard, not a W3C one, which matters mainly because the actual document is short and worth reading.
The rule people get wrong is group selection. A crawler matches its own product token against the User-agent lines, obeys the group it matched, and falls back to User-agent: * only when nothing matched its name. It is not a cascade.
User-agent: *
Disallow: /admin/
Disallow: /cart
User-agent: Googlebot
Disallow: /drafts/Googlebot will happily crawl /admin/ and /cart here. It found a group carrying its own name, so the * group is not addressed to it at all — that block is the fallback for everyone else. If you want Googlebot bound by the shared rules as well, repeat them under its heading. The robots.txt generator keeps directives grouped per user agent rather than concatenating them, because a rule under the wrong heading fails silently and looks fine in the file.
How a path is actually matched
Rule paths are prefix matches against the URL path. They are not filenames, and they are not shell globs.
Disallow: /adminThat blocks /admin/, but it also blocks /administrator, /admin-guide.html and /admins/list. Write /admin/ with the trailing slash if you mean the directory only. Paths are case-sensitive; directive names and user-agent tokens are not.
There are exactly two metacharacters. * matches any run of characters, and a trailing $ anchors the match to the end of the URL. So /*.pdf$ blocks every URL ending in .pdf (but not /report.pdf?download=1, because Google matches a rule against the path and the query string together), and /private$ blocks /private while leaving /private-beta alone.
When an Allow and a Disallow both match, the more specific one wins, and specificity is measured in characters of the rule path.
User-agent: *
Disallow: /blog
Allow: /blog/publicFor /blog/public/post-1, Disallow: /blog matches at 5 characters and Allow: /blog/public matches at 12. The longer rule wins and the URL is crawled. For /blog/drafts/post-2, only the Disallow matches, so it is blocked. On an exact tie the Allow wins:
Allow: /reports
Disallow: /reportsBoth are 8 characters, so /reports is crawlable. Piling on more Disallow lines will not change that; deleting the Allow will.
That length-then-Allow resolution is Google's behaviour and what RFC 9309 specifies. Older crawlers built to the 1994 convention take the first matching rule in file order instead. You can satisfy both by putting Allow lines above the Disallow lines they carve exceptions out of, which is the order the generator on this site emits.
A real file, and what it is actually for
Here is the whole of this site's robots.txt:
User-Agent: Googlebot
User-Agent: Bingbot
User-Agent: OAI-SearchBot
User-Agent: ChatGPT-User
User-Agent: Claude-SearchBot
User-Agent: Claude-User
User-Agent: PerplexityBot
User-Agent: Perplexity-User
User-Agent: Applebot
Allow: /
Disallow: /*__next.*.txt
Disallow: /*index.txt
User-Agent: *
Allow: /
Disallow: /*__next.*.txt
Disallow: /*index.txt
Sitemap: https://toolsbay.co/sitemap.xmlThe same two Disallow lines in each group, and neither of them is hiding anything. The first group names the search and answer crawlers outright. It has to repeat the rules, because a crawler that finds a group naming it follows that group and ignores * entirely. The Next.js App Router writes a plain-text prefetch payload for every route segment so the client-side router can fetch a page ahead of a click — 455 files for 90 pages when this was written, served as text/plain with a 200, each one embedding the copy of a real page.
Crawlers took them. In our access logs for the week to 23 August 2026, seven of the twelve busiest paths on the entire site were these payloads, and /tools/pdf-to-word/__next._tree.txt was fetched 383 times against 313 for /tools/pdf-to-word/ itself. The site was six days old. Google had barely finished discovering the real pages.
Disallow is the right instrument here precisely because the problem is crawling and not indexing. A noindex header would have kept them out of results too, and it would have cost a fetch each: the crawler downloads every one of those files and only then throws them away. Disallow prevents the request. That is the whole distinction, in the one case where it cuts the other way.
Both patterns used to end in $, and that was a mistake worth describing. The router never requests /tools/merge-pdf/index.txt; it requests /tools/merge-pdf/index.txt?_rsc= followed by a short hash. Because Google matches a rule against the path and the query string together, the anchored rules matched the files on disk and missed every request a crawler actually followed. Without the $, each pattern matches any URL that contains those characters, and a rule that broad at a site's root is one typo away from Disallow: /. So there is a test beside the file asserting both directions — that all eleven payload shapes match, with and without a query string, and that every URL in the sitemap, plus /robots.txt and /sitemap.xml, does not. A Disallow that quietly matches too much does not throw an error. It removes your traffic several weeks later.
There is no Host: line. The framework will write one if asked, but Google ignores it; only Yandex ever used it, to nominate a canonical mirror. So it came out.
Directives that do not do what their name suggests
Crawl-delay is ignored by Googlebot outright. Bing and Yandex honour it; Google's documentation lists it as unsupported and points you at the crawl rate setting in Search Console instead. Leaving it in is not an error, but do not expect it to slow Google down.
Noindex: as a robots.txt directive was never in the standard. Google supported it undocumented for years and switched it off on 1 September 2019, along with nofollow and crawl-delay in the same file. Any tutorial still recommending it predates that.
The file is public, and that is the point
Anyone can read https://yoursite.com/robots.txt. A crawler that ignores it faces no obstacle — there is no enforcement, only convention. So Disallow: /private.pdf is not a way to protect a PDF. It is a way to publish the exact path of a PDF you would rather people did not read, to every scraper on the internet, in a file they are guaranteed to fetch first. The same goes for Disallow: /secret-admin. If something needs to be private, it needs authentication.
Placement and failure modes
The file lives at the root of the origin and nowhere else. It is matched by scheme, host and port, so https://example.com/robots.txt says nothing about http://example.com/, and blog.example.com needs its own. A robots.txt in a subdirectory is read by nothing.
Google parses the first 500 kibibytes and ignores the rest, and generally caches the file for up to 24 hours, so an edit is not instant. The response code matters more than most people realise: a 404 means no restrictions and Google crawls everything, while a 5xx is treated as the whole site is disallowed for as long as the error persists. A robots.txt behind a flaky server is a worse outcome than no robots.txt at all.
The Sitemap: line belongs to no user-agent group and can appear anywhere in the file — first line, last line, between groups. It must be an absolute URL, and you can list more than one. Generate the file it points at with the XML sitemap generator, then confirm the URL resolves before you rely on it.
Everything in here is a request. It is a good one, and well-behaved crawlers honour it. It is not a lock and it is not an index control. Used for what it is — steering finite crawl budget away from things that are not worth fetching — it does real work.