Two different instructions, often confused
A robots.txt disallow rule and a noindex directive answer different questions. Disallow tells a compliant crawler not to fetch a URL at all. Noindex, delivered via a meta tag or X-Robots-Tag header, tells a compliant crawler it may fetch the page but should not include it in an index or index-derived answer surface. Google’s own robots meta tag documentation describes this distinction for its own search index; the same conceptual split is the reasonable mental model for AI crawlers, though each AI crawler operator’s actual compliance with these directives is governed by its own stated policy, not Google’s.
The classic mistake: disallow and noindex together
If a URL is disallowed in robots.txt, a compliant crawler will never fetch it, which means it will never see a noindex directive placed on that page, because the directive lives on the page itself. The noindex instruction becomes unreachable and ineffective. If the goal is genuinely to keep a URL out of an index while still allowing it to be crawled for other reasons, use noindex alone, without a matching disallow rule.
When to use disallow
Use disallow for URLs that should not be fetched at all: low-value parameterized duplicates, internal search result pages, staging paths that should never be accessible, or paths that are expensive to serve to crawlers at crawl-worthy volume. Disallow reduces crawl load on top of affecting indexing.
When to use noindex
Use noindex for pages that should remain crawlable, perhaps because internal links or other signals still need to be followed through them, but should not appear in an index or be used as a direct citable answer: thin utility pages, duplicate content kept for a legitimate reason, or draft content intentionally excluded pending review. This site’s own guide drafts use exactly this pattern pending owner approval.
Neither is enforcement
Both are stated, voluntary directives. A crawler that ignores robots.txt entirely will also ignore a noindex directive; neither mechanism can force a non-compliant client to comply. For genuine access control, a server-side block, authentication, or edge-level bot rule is required, not a crawling or indexing directive.
See robots txt vs noindex, robots policy is not enforcement, and staging site indexing prevention.
Frequently asked questions
Can I use both disallow and noindex on the same URL?
Not effectively together. Disallow prevents the crawler from ever seeing the noindex instruction on that page, so combining them on the same URL to achieve exclusion is self-defeating; pick one based on whether the page should be fetched at all.
Does noindex stop an AI crawler from training on a page?
Noindex and robots.txt disallow rules address crawling and indexing. Training-data use is typically governed by a separate directive some crawlers publish distinctly from their search-access directive; check each crawler own documented policy rather than assuming one directive covers both.