Robots.txt causes more disproportionate damage than any other configuration file in technical SEO — one wrong disallow rule can start pages dropping out of Google within days. I run a robots.txt audit in the first hour of every technical engagement. The errors I find range from legacy rules blocking URL structures that no longer exist, to staging directives accidentally deployed to production, to conflicting rules added by different developers over the years. This guide covers the syntax, the strategy, and the mistakes to guard against.
Robots.txt Syntax — The Complete Reference
# Robots.txt syntax reference # Location: https://yourdomain.com/robots.txt # Allow all crawlers everything (same as having no robots.txt) User-agent: * Allow: / # Block all crawlers from everything (DANGEROUS on a live site) User-agent: * Disallow: / # Block Googlebot from the /admin/ directory User-agent: Googlebot Disallow: /admin/ # Block all crawlers from PDF files User-agent: * Disallow: /*.pdf$ # Allow a path inside a blocked directory User-agent: * Disallow: /private/ Allow: /private/public-section/ # Tell crawlers where your sitemap lives Sitemap: https://yourdomain.com/sitemap.xml
The Critical Rules to Understand Before Touching Robots.txt
Disallow Does Not Mean Noindex
This is the most misunderstood fact about robots.txt. Disallow stops Googlebot crawling a URL — it doesn't stop Google indexing it. If other pages link to a disallowed URL, Google can still index the URL without its content, and it may appear in results with no description. To keep a page out of the index, use a noindex meta tag or X-Robots-Tag header — and leave the page crawlable so Google can see that tag. Disallowing and noindexing the same URL doesn't work, because Google never sees the noindex.
Robots.txt Is a Courtesy, Not a Lock
Reputable crawlers — Googlebot, Bingbot and other major bots — follow your directives. Malicious bots and scrapers ignore them. Robots.txt is also publicly readable, so listing a sensitive path in it actually advertises it. Never rely on robots.txt for privacy or security; use authentication and server-side access controls for anything genuinely private.
The Most Specific Rule Wins
For Googlebot, when Allow and Disallow rules conflict, the rule with the longest matching path wins, regardless of order. Disallow: /private/ plus Allow: /private/public/ correctly allows the public subfolder. If two matching rules are equally specific, Google uses the least restrictive one (Allow).
Wildcards Work for Google and Bing — Not Every Crawler
Googlebot supports * (any sequence of characters) and $ (end of URL). Disallow: /*? blocks every URL with a query string; Disallow: /*.pdf$ blocks PDFs. Not all crawlers support wildcards, so check a crawler's documentation before relying on them for bots other than Googlebot and Bingbot.
What You Should Block — and What You Should Never Block
| URL Type | Block in Robots.txt? | Why |
|---|---|---|
| Admin areas (/wp-admin/) | Usually — Disallow | No search value and wastes crawling. On WordPress keep Allow: /wp-admin/admin-ajax.php, which front-end features use. It gives no security benefit — robots.txt is public. |
| Staging or development sites | Yes — on staging only | Keeps duplicate staging copies out of crawling; password protection is safer still. Never let the staging file reach production. |
| Internal search results (?s=, /search?q=) | Yes — Disallow | Creates near-infinite crawl traps of thin, duplicate pages |
| Cart and checkout pages | Yes — Disallow | No ranking value; session parameters create duplicate URLs |
| Tag and category archives (WordPress) | Sometimes — evaluate | Block if they're thin duplicates; allow if they're genuine landing pages with unique value |
| Core content pages (blog, products, services) | Never | These are the pages you need indexed and ranked |
| CSS and JavaScript files | Never | Google renders pages with your CSS and JavaScript; blocking them can break how Google sees your content and layout |
| XML sitemaps | Never | Google needs to fetch the sitemap to use it |
| Images | Only if deliberate | Blocking Googlebot-Image removes you from image search — do it only if you truly don't want image traffic |
Robots.txt for AI Crawlers — The 2026 Addition
AI companies now run separate user agents for model training and for AI search and citations, and blocking one doesn't block the other. The main directives:
# OpenAI training crawler User-agent: GPTBot Disallow: / # OpenAI search crawler — blocking it removes you from ChatGPT search results User-agent: OAI-SearchBot Disallow: / # Google-Extended — a control token, not a separate crawler. # Opts content out of Gemini model training/grounding; # does NOT affect Google Search or AI Overviews. User-agent: Google-Extended Disallow: / # Anthropic training crawler (Claude-SearchBot and Claude-User are separate) User-agent: ClaudeBot Disallow: / # Perplexity search crawler — blocking it removes you from Perplexity answers User-agent: PerplexityBot Disallow: / # To stay citable but opt out of training: # block GPTBot / ClaudeBot / Google-Extended, allow OAI-SearchBot and PerplexityBot
GPTBot (OpenAI's training crawler) and OAI-SearchBot (the crawler behind ChatGPT search) are different user agents with different consequences. Blocking GPTBot keeps your content out of future model training. Blocking OAI-SearchBot keeps you out of ChatGPT search answers and citations. If you want to appear in AI search answers but not contribute training data, block the training agents and allow the search agents. Most "block all AI bots" decisions are made without separating those two outcomes. If AI visibility matters to you, also see our guide to llms.txt and AI visibility.
"The robots.txt error I encounter most often — and find most alarming — is staging directives accidentally deployed to production. A team protects staging with User-agent: * and Disallow: /, which is fine for staging. Then a deployment pushes that file to the live site. Within a day or two Googlebot stops being able to recrawl pages, and as the days pass rankings slip because Google can't refresh anything. By the time anyone spots the cause, the site has lost weeks of crawling. For every client I onboard, I add robots.txt to the deployment checklist with one specific step: diff the robots.txt being deployed against the last confirmed production version before anything goes live. One line item prevents the most catastrophic avoidable SEO failure in the developer workflow."
Robots.txt works together with your XML sitemap and is one of your main levers for managing crawl budget.
Frequently Asked Questions
Sitemap: https://yourdomain.com/sitemap.xml costs nothing and helps every crawler that reads robots.txt — Bing and AI crawlers as well as Google — find your sitemap. Use the full absolute URL, and still submit the sitemap in Search Console so you can monitor it.The Bottom Line
Robots.txt is a small file with outsized power to help or harm your SEO. Block internal search results, cart pages, admin areas and staging environments; never block core content, CSS, JavaScript or your sitemap. Remember that Disallow stops crawling but not indexing — use noindex on crawlable pages to keep them out of search. Treat AI crawlers deliberately, separating training agents (GPTBot, ClaudeBot, Google-Extended) from search agents (OAI-SearchBot, PerplexityBot). Reference your sitemap, check changes with Search Console's robots.txt report and URL Inspection, and put a robots.txt diff check in your deployment process. One wrong line can take a site out of Google; one careful check prevents it.
Sources: Google Search Central — Robots.txt introduction; Google — How Google interprets robots.txt; Google — Common crawlers and Google-Extended; OpenAI — Crawlers overview.
Driven by advanced SEO expertise, deep marketing analytics, high-impact content strategy
With 5+ years of hands-on experience, I specialize in holistic search strategies that don’t just rank—they drive real, measurable business growth. I’ve worked across industries including healthcare, hospitality, legal, e-commerce, and professional services, helping brands dominate their target markets. My approach bridges the gap between raw data and creative execution. Every strategy I build is rooted in rigorous market analysis, structured SEO frameworks, and tailored content ecosystems—no templates, no shortcuts. Whether you’re a single-location brand or scaling across multiple cities, I create data-driven marketing systems designed to compound results and grow with you.
Not Sure What Your Robots.txt Is Blocking?
Get a technical SEO audit that checks your robots.txt, AI crawler access and indexation — before a bad rule costs you rankings.
Book a Free SEO Audit →