Robots.txt URL Tester
See which crawlers may fetch a URL — and the exact rule that decides it.
robots.txt to test
The rules are read in your browser and never uploaded. Results update as you type.
Fetching robots.txt from MySmartCoPilot’s server…
URLs and crawlers
Leave it empty to test the home page (/). Paths are compared exactly as RFC 9309 describes, including the query string.
Results
Select a result to see the rule that decided it.
The file
Syntax checks
Groups
Lines
About the Robots.txt URL Tester
Enter a website and the tester fetches its robots.txt, or paste rules you are still writing. Then test up to 100 URLs at a time against the crawlers you care about — Googlebot and Bingbot, and the AI crawlers GPTBot, ClaudeBot, Google-Extended, CCBot and PerplexityBot — and see for each one whether it may fetch the URL, and the exact line that decides it.
Matching follows RFC 9309, the Robots Exclusion Protocol standard: groups for the same crawler are combined, the longest matching rule wins, Allow wins a tie, and * and $ work as wildcards. The file view lists every Sitemap: line and flags syntax problems such as misspelt fields, rules outside a group and rules that can never match.
How to use it
- Choose Fetch from a website and enter a domain or page URL, or choose Paste rules and paste a robots.txt.
- List the URLs or paths to test, one per line — or leave the box empty to test the home page.
- Tick the crawlers to test, and type any others (product tokens such as
DuckDuckBot) in Other crawlers. - Press Test URLs. Each cell says Allowed, Blocked or Unclear; select it to see the deciding rule, the group used and every rule that matched, highlighted in the file.
- Download the results as CSV, or copy them into a ticket.
Examples
User-agent: * Disallow: /admin/ User-agent: GPTBot User-agent: CCBot Disallow: /
/ → Googlebot Allowed · GPTBot Blocked (line 6) · CCBot Blocked (line 6) · PerplexityBot Allowed /admin/users → Blocked for all (line 2, or line 6 for GPTBot and CCBot)
GPTBot and CCBot have their own group, so the “*” rules do not apply to them at all — a crawler follows only the most specific group that names it.
User-agent: * Disallow: /shop/ Allow: /shop/sale
/shop/sale/1 → Allowed by “Allow: /shop/sale” (10 characters) over “Disallow: /shop/” (6) /shop/cart → Blocked by “Disallow: /shop/”
Common uses
- Checking that a new robots.txt does not block pages you want in Google before you deploy it.
- Finding out whether a site blocks AI crawlers such as GPTBot, ClaudeBot or CCBot.
- Debugging why a crawler skips a section: the deciding line is highlighted.
- Auditing which sitemaps a site announces in its robots.txt.
How a crawler reads robots.txt (RFC 9309)
- A crawler looks for groups whose User-agent line names its product token, case-insensitively; several such groups are combined. Only if none exists does it use the
User-agent: *group, and with neither, no rules apply. - Within the group, the most specific (longest) matching rule decides. When an Allow and a Disallow rule are equally long, Allow wins.
*matches any run of characters and$at the end anchors the end of the URL;#starts a comment. Paths are case-sensitive and compared from the first character, query string included.- Paths are compared percent-encoded:
/foo/bar/%62%61%7Acounts as/foo/bar/baz, and a literal*in a URL is matched by%2Ain a rule. /robots.txtitself is always allowed.
See RFC 9309 §2.2 and Google’s “How Google interprets the robots.txt specification”. Two Google-only extras are applied to Google’s crawlers alone: some of them have a second token (“you need to match only one crawler token for a rule to apply”), so Googlebot-Image, -News and -Video, Google-InspectionTool and Google-CloudVertexBot use the Googlebot group when no group names them, and GoogleOther-Image and -Video the GoogleOther group; and the special-case crawlers AdsBot-Google, AdsBot-Google-Mobile, Mediapartners-Google and APIs-Google ignore *.
When robots.txt itself cannot be read
- 4xx (for example 404): the file is “unavailable” and crawlers may crawl everything. Google treats all 4xx answers this way except 429, and asks site owners not to use 401 or 403 to limit crawling.
- More than five redirects: RFC 9309 lets crawlers treat the file as unavailable; Google treats it as a 404.
- 5xx, timeouts and network errors: the file is “unreachable”, and RFC 9309 says crawlers must assume complete disallow. Google stops crawling the site for the first 12 hours, then uses its last good copy for up to 30 days.
- Crawlers read at least the first 500 KiB (Google exactly 500 KiB) and may cache the file for up to 24 hours.
The tester shows which of these applies and applies it to every crawler and URL.
AI crawlers and their tokens
Each operator documents its own tokens:
- GPTBot (OpenAI) crawls content that may be used to train its models; OAI-SearchBot is for ChatGPT search. ChatGPT-User fetches pages for a user, and OpenAI says robots.txt rules may not apply to it. OpenAI
- ClaudeBot (Anthropic) is the training crawler; Claude-SearchBot and Claude-User are separate, and Anthropic says all three honour robots.txt. Anthropic
- Google-Extended is not a crawler but a token that controls use of Google-crawled content for Gemini training and grounding; it does not affect Google Search. Google
- CCBot builds Common Crawl’s free, open archive of web pages, which anyone can download and use. Common Crawl
- PerplexityBot surfaces sites in Perplexity’s search results; Perplexity-User fetches pages for a user and, Perplexity says, generally ignores robots.txt. Perplexity
To block these, add a group for them with Disallow: / — the Robots.txt Generator has a preset.
Limitations
- robots.txt is a request, not access control: well-behaved crawlers follow it, others ignore it. Protect private content with authentication.
- Blocking a URL does not keep it out of search results: Google can still list it (without a description) if other pages link to it. Use a noindex robots meta tag on a crawlable page instead — the Indexability Checker shows both.
- MySmartCoPilot fetches robots.txt as MySmartCoPilotBot. A site that serves different files to different crawlers (rare) may give Googlebot other rules.
- Crawlers that fetch on behalf of a user (such as ChatGPT-User or Perplexity-User) may not follow robots.txt at all.
- Each run tests up to 100 URLs; the file view shows the first 3,000 lines, although every line is tested.
Privacy
In “Fetch from a website” mode, the address you enter is sent to MySmartCoPilot’s server, which requests only that site’s /robots.txt and returns it; nothing is stored, and the log keeps only the host name, status and time. Pasted rules and the URLs you test never leave your browser.
Frequently asked questions
Why is a URL allowed although a Disallow rule matches it?
Because a longer Allow rule also matches: the most specific rule wins, and when an Allow and a Disallow rule are equally long, Allow wins (RFC 9309 §2.2.2). Select the cell to see every matching rule and its length.
Why does a crawler ignore my “User-agent: *” rules?
A crawler that has its own group follows only that group. If you add User-agent: GPTBot with one rule, GPTBot no longer reads the * group — repeat the rules it should also follow in its own group.
Why is everything blocked although my robots.txt is empty?
Check the status line above the results. If the server answered with a 5xx error or did not answer, RFC 9309 tells crawlers to assume everything is disallowed until the file can be read again. Make /robots.txt return 200 with your rules, or 404 if you have none.
Does robots.txt on example.com cover www.example.com or blog.example.com?
No. Each protocol, host and port has its own robots.txt: https://www.example.com/robots.txt covers only https://www.example.com/. The tester marks URLs on another host.
How do I block AI crawlers but stay in Google?
Give GPTBot, ClaudeBot, CCBot, Google-Extended and the other training tokens their own group with Disallow: /, and leave Googlebot and Bingbot allowed. Blocking Google-Extended does not affect Google Search. Then test the file here with those crawlers ticked.