Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

Robots.txt URL Tester

See which crawlers may fetch a URL — and the exact rule that decides it.

Network Uses live data Free, no sign-up

robots.txt to test

MySmartCoPilot’s server fetches /robots.txt of that host (the first 500 KiB, following up to five redirects). A robots.txt only covers its own protocol, host and port: www.example.com and example.com each have their own.

URLs and crawlers

Leave it empty to test the home page (/). Paths are compared exactly as RFC 9309 describes, including the query string.

Crawlers

Next steps

About the Robots.txt URL Tester

Enter a website and the tester fetches its robots.txt, or paste rules you are still writing. Then test up to 100 URLs at a time against the crawlers you care about — Googlebot and Bingbot, and the AI crawlers GPTBot, ClaudeBot, Google-Extended, CCBot and PerplexityBot — and see for each one whether it may fetch the URL, and the exact line that decides it.

Matching follows RFC 9309, the Robots Exclusion Protocol standard: groups for the same crawler are combined, the longest matching rule wins, Allow wins a tie, and * and $ work as wildcards. The file view lists every Sitemap: line and flags syntax problems such as misspelt fields, rules outside a group and rules that can never match.

How to use it

  1. Choose Fetch from a website and enter a domain or page URL, or choose Paste rules and paste a robots.txt.
  2. List the URLs or paths to test, one per line — or leave the box empty to test the home page.
  3. Tick the crawlers to test, and type any others (product tokens such as DuckDuckBot) in Other crawlers.
  4. Press Test URLs. Each cell says Allowed, Blocked or Unclear; select it to see the deciding rule, the group used and every rule that matched, highlighted in the file.
  5. Download the results as CSV, or copy them into a ticket.

Examples

Blocking AI training but not search
Input
User-agent: *
Disallow: /admin/

User-agent: GPTBot
User-agent: CCBot
Disallow: /
Result
/ → Googlebot Allowed · GPTBot Blocked (line 6) · CCBot Blocked (line 6) · PerplexityBot Allowed
/admin/users → Blocked for all (line 2, or line 6 for GPTBot and CCBot)

GPTBot and CCBot have their own group, so the “*” rules do not apply to them at all — a crawler follows only the most specific group that names it.

The longest rule wins
Input
User-agent: *
Disallow: /shop/
Allow: /shop/sale
Result
/shop/sale/1 → Allowed by “Allow: /shop/sale” (10 characters) over “Disallow: /shop/” (6)
/shop/cart → Blocked by “Disallow: /shop/”

Common uses

  • Checking that a new robots.txt does not block pages you want in Google before you deploy it.
  • Finding out whether a site blocks AI crawlers such as GPTBot, ClaudeBot or CCBot.
  • Debugging why a crawler skips a section: the deciding line is highlighted.
  • Auditing which sitemaps a site announces in its robots.txt.

How a crawler reads robots.txt (RFC 9309)

  • A crawler looks for groups whose User-agent line names its product token, case-insensitively; several such groups are combined. Only if none exists does it use the User-agent: * group, and with neither, no rules apply.
  • Within the group, the most specific (longest) matching rule decides. When an Allow and a Disallow rule are equally long, Allow wins.
  • * matches any run of characters and $ at the end anchors the end of the URL; # starts a comment. Paths are case-sensitive and compared from the first character, query string included.
  • Paths are compared percent-encoded: /foo/bar/%62%61%7A counts as /foo/bar/baz, and a literal * in a URL is matched by %2A in a rule.
  • /robots.txt itself is always allowed.

See RFC 9309 §2.2 and Google’s “How Google interprets the robots.txt specification”. Two Google-only extras are applied to Google’s crawlers alone: some of them have a second token (“you need to match only one crawler token for a rule to apply”), so Googlebot-Image, -News and -Video, Google-InspectionTool and Google-CloudVertexBot use the Googlebot group when no group names them, and GoogleOther-Image and -Video the GoogleOther group; and the special-case crawlers AdsBot-Google, AdsBot-Google-Mobile, Mediapartners-Google and APIs-Google ignore *.

When robots.txt itself cannot be read

  • 4xx (for example 404): the file is “unavailable” and crawlers may crawl everything. Google treats all 4xx answers this way except 429, and asks site owners not to use 401 or 403 to limit crawling.
  • More than five redirects: RFC 9309 lets crawlers treat the file as unavailable; Google treats it as a 404.
  • 5xx, timeouts and network errors: the file is “unreachable”, and RFC 9309 says crawlers must assume complete disallow. Google stops crawling the site for the first 12 hours, then uses its last good copy for up to 30 days.
  • Crawlers read at least the first 500 KiB (Google exactly 500 KiB) and may cache the file for up to 24 hours.

The tester shows which of these applies and applies it to every crawler and URL.

AI crawlers and their tokens

Each operator documents its own tokens:

  • GPTBot (OpenAI) crawls content that may be used to train its models; OAI-SearchBot is for ChatGPT search. ChatGPT-User fetches pages for a user, and OpenAI says robots.txt rules may not apply to it. OpenAI
  • ClaudeBot (Anthropic) is the training crawler; Claude-SearchBot and Claude-User are separate, and Anthropic says all three honour robots.txt. Anthropic
  • Google-Extended is not a crawler but a token that controls use of Google-crawled content for Gemini training and grounding; it does not affect Google Search. Google
  • CCBot builds Common Crawl’s free, open archive of web pages, which anyone can download and use. Common Crawl
  • PerplexityBot surfaces sites in Perplexity’s search results; Perplexity-User fetches pages for a user and, Perplexity says, generally ignores robots.txt. Perplexity

To block these, add a group for them with Disallow: / — the Robots.txt Generator has a preset.

Limitations

  • robots.txt is a request, not access control: well-behaved crawlers follow it, others ignore it. Protect private content with authentication.
  • Blocking a URL does not keep it out of search results: Google can still list it (without a description) if other pages link to it. Use a noindex robots meta tag on a crawlable page instead — the Indexability Checker shows both.
  • MySmartCoPilot fetches robots.txt as MySmartCoPilotBot. A site that serves different files to different crawlers (rare) may give Googlebot other rules.
  • Crawlers that fetch on behalf of a user (such as ChatGPT-User or Perplexity-User) may not follow robots.txt at all.
  • Each run tests up to 100 URLs; the file view shows the first 3,000 lines, although every line is tested.

Privacy

In “Fetch from a website” mode, the address you enter is sent to MySmartCoPilot’s server, which requests only that site’s /robots.txt and returns it; nothing is stored, and the log keeps only the host name, status and time. Pasted rules and the URLs you test never leave your browser.

Frequently asked questions

Why is a URL allowed although a Disallow rule matches it?

Because a longer Allow rule also matches: the most specific rule wins, and when an Allow and a Disallow rule are equally long, Allow wins (RFC 9309 §2.2.2). Select the cell to see every matching rule and its length.

Why does a crawler ignore my “User-agent: *” rules?

A crawler that has its own group follows only that group. If you add User-agent: GPTBot with one rule, GPTBot no longer reads the * group — repeat the rules it should also follow in its own group.

Why is everything blocked although my robots.txt is empty?

Check the status line above the results. If the server answered with a 5xx error or did not answer, RFC 9309 tells crawlers to assume everything is disallowed until the file can be read again. Make /robots.txt return 200 with your rules, or 404 if you have none.

Does robots.txt on example.com cover www.example.com or blog.example.com?

No. Each protocol, host and port has its own robots.txt: https://www.example.com/robots.txt covers only https://www.example.com/. The tester marks URLs on another host.

How do I block AI crawlers but stay in Google?

Give GPTBot, ClaudeBot, CCBot, Google-Extended and the other training tokens their own group with Disallow: /, and leave Googlebot and Bingbot allowed. Blocking Google-Extended does not affect Google Search. Then test the file here with those crawlers ticked.

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.