Robots.txt is one of the oldest and most fundamental tools in SEO, yet it is routinely misconfigured even on major sites. While the basic syntax is simple — allow and disallow directives for user agents — the advanced patterns for wildcard matching, crawl-delay, sitemap declarations, and section-level control can make the difference between efficient crawl budget utilization and catastrophic indexing failures.
Robots.txt Fundamentals
The robots.txt file lives at the root of your domain and provides crawl instructions to search engine bots. It is a suggestion, not a command — well-behaved bots follow it, but malicious bots ignore it. Google's crawlers follow robots.txt directives strictly. The file uses plain text with User-agent, Disallow, Allow, and Sitemap directives. Rules are processed in order of specificity, not top-to-bottom.
Robots.txt blocks crawling, not indexing. If a page is disallowed in robots.txt but has external links pointing to it, Google may still index the URL — showing it in search results with no snippet because it cannot crawl the content. Use noindex meta tags or X-Robots-Tag headers to prevent indexing.
Wildcard Patterns
Google supports two wildcard characters in robots.txt: the asterisk for matching any sequence of characters and the dollar sign for matching the end of a URL. These wildcards enable powerful pattern matching that can target specific URL structures without listing every individual URL.
- Disallow: /*?sort= — blocks all URLs containing the sort parameter
- Disallow: /*.pdf$ — blocks all URLs ending in .pdf
- Disallow: /category/*/page/ — blocks pagination within category paths
- Allow: /category/$ — allows the exact /category/ page while other rules block deeper URLs
- Disallow: /*&color=* — blocks URLs with the color filter parameter
Precedence and Specificity
When multiple rules could apply to a URL, Google uses the most specific matching rule — the rule with the longest matching path. This means you can create broad disallow rules and then create more specific allow rules to grant access to individual URLs or patterns within the blocked section. Understanding precedence is critical for complex robots.txt configurations.
Example: Blocking Filters While Allowing Categories
To block faceted navigation parameters while keeping category pages crawlable, use a combination of broad disallow patterns for filter parameters with specific allow rules for clean category URLs. The allow rule for the exact category path will take precedence over the broader disallow rule for parameterized variations.
User-Agent Specific Rules
Different search engines use different user agents. You can create rules that apply to all bots using User-agent: * or target specific bots like Googlebot, Bingbot, or GPTBot individually. This is useful when you want to allow Google to crawl certain content but block AI training crawlers, or when you need different crawl rules for different search engines.
Common Robots.txt Mistakes
- Blocking CSS and JavaScript files: prevents Google from rendering your pages, hurting rankings
- Blocking the entire site accidentally: a single misplaced Disallow: / blocks everything
- Using robots.txt to hide pages from the index: it blocks crawling but not indexing
- Forgetting the trailing slash: Disallow: /admin blocks /admin, /admin/, and /administrator
- Conflicting rules without understanding precedence: more specific rules always win
- Not including a Sitemap directive: always point to your XML sitemap in robots.txt
We once audited a site that had accidentally blocked their entire /blog/ directory in robots.txt during a server migration. It went unnoticed for four months. They lost 60 percent of their organic blog traffic before the error was discovered. Always audit your robots.txt after any server or infrastructure changes.
Testing and Monitoring
Use Google Search Console's robots.txt tester to validate your rules before deploying changes. Test specific URLs to see whether they are blocked or allowed by your current configuration. After deploying changes, monitor Search Console's crawl stats for any unexpected drops in crawl activity. Set up change monitoring for your robots.txt file to alert you if it is modified unexpectedly — this can happen during CMS updates or server configuration changes.