Robots.txt Generator & Tester
Create, check and test robots.txt rules exactly the way Googlebot reads them.
Start from a preset
Presets change the rules for all crawlers (*); other groups stay as they are.
Block AI training crawlers
Only crawler tokens confirmed in each operator’s own documentation (linked below). Search crawlers (Googlebot, Bingbot, OAI-SearchBot…) are not affected.
-
Crawls content that may be used to train OpenAI’s generative AI foundation models. OAI-SearchBot (ChatGPT search) is a separate token and stays allowed. OpenAI documentation
-
Collects web content that could contribute to Anthropic’s model training. Claude-SearchBot and Claude-User are separate tokens. Anthropic documentation
-
Controls whether content Google crawls may be used to train future Gemini models and for grounding. It is not a crawler and does not affect Google Search. Google documentation
-
Controls whether content crawled by Applebot may be used to train Apple’s foundation models. It does not crawl and does not remove you from Apple’s search features. Apple documentation
-
Builds Common Crawl’s free, open archive of web pages, which anyone may download and reuse — including to train AI models. Common Crawl documentation
-
Crawls the web for uses such as training foundation AI models. Link previews (facebookexternalhit) are not affected. Meta documentation
-
Used to improve Amazon’s products and services; content may be used to train Amazon AI models. Amzn-SearchBot (search features such as Alexa) is separate and not used for training. Amazon documentation
Rules by crawler
Sitemaps
Full URLs, e.g. https://example.com/sitemap.xml. Any crawler that reads robots.txt can find them here.
Paste the contents of a robots.txt file. Open https://your-site/robots.txt in a new tab to copy it.
Test a URL
| Line | Rule | Length | Matches? |
|---|
Validation
No problems found. Every line is valid for Google.
How crawlers see this file
| Line | User agents | Rules | Crawl-delay |
|---|
About the Robots.txt Generator & Tester
A robots.txt file tells crawlers which parts of your site they may fetch. This tool builds one from a simple form — groups of rules per crawler, sitemap lines and one-click presets — and checks it as you go.
Paste an existing file to validate it line by line, then test any URL against it. The tester implements Google’s documented rules: it picks the group for the crawler (falling back to *), compares rules by length so the most specific one wins, lets Allow win a tie, supports the * and $ wildcards and reads only the first 500 KiB, like Googlebot. It tells you which line decided the result and why.
How to use it
- Start from a preset — Allow all, Block all, Block admin areas — or edit the
*group directly. - Add rules: choose Disallow or Allow and type a path such as
/private/or/*.pdf$. Add more groups for specific crawlers if they need different rules. - Optionally tick the AI training crawlers you want to opt out of and press Add to robots.txt.
- Add the full URL of your sitemap, then copy or download the file and upload it to the root of your site, e.g.
https://example.com/robots.txt. - Switch to Test & validate to paste any robots.txt and check whether a URL is allowed for Googlebot, GPTBot or any other crawler.
Examples
User-agent: * Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://example.com/sitemap.xml
/wp-admin/admin-ajax.php is allowed because the Allow rule is longer (more specific) than the Disallow rule.
User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: CCBot Disallow: /
Googlebot, Bingbot and the AI companies’ separate search crawlers (such as OAI-SearchBot) are not named, so they keep crawling.
How Google reads robots.txt
- One group per crawler. A crawler follows the group with the most specific matching
User-agentline (case-insensitive); all other groups are ignored. If several groups name the same crawler, they are combined. Only if no group names it does it followUser-agent: *— except Google’s special-case crawlers AdsBot-Google, AdsBot-Google-Mobile, Mediapartners-Google and APIs-Google, which ignore*and obey only a group that names them. - Longest rule wins. Among the
AllowandDisallowrules that match a URL, the one with the longest path wins. If anAllowand aDisalloware equally long, Google uses the least restrictive one —Allow. - Wildcards.
*matches any sequence of characters;$marks the end of the URL./*.pdf$matches/files/report.pdfbut not/report.pdf?download=1. - Prefix matching, case-sensitive.
Disallow: /fishalso blocks/fish.htmland/fishheads, but not/Fish.asp. - Location and size. The file must be at the root of each host (
https://example.com/robots.txt); it applies only to that protocol, host and port. Google reads the first 500 KiB, caches the file for up to about 24 hours and treats a 404 as “no restrictions”. - Only four fields. Google supports
user-agent,allow,disallowandsitemap.Crawl-delay,Host,Noindexand other fields are ignored.
Source: Google’s robots.txt specification and RFC 9309.
Blocking AI training crawlers
The AI preset only lists crawler tokens confirmed in each operator’s own documentation:
GPTBot— OpenAI: crawls content that may be used to train its generative AI models. OAI-SearchBot, used for ChatGPT search, is separate.ClaudeBot— Anthropic: collects content that could contribute to model training. Claude-SearchBot and Claude-User are separate.Google-Extended— Google: controls use for Gemini training and grounding; it does not affect Google Search.Applebot-Extended— Apple: controls use for training Apple’s foundation models; it does not crawl or affect Apple search.CCBot— Common Crawl: builds a free, open web archive that anyone may download and reuse, AI developers included.meta-externalagent— Meta: crawls for uses such as training foundation models.Amazonbot— Amazon: content may be used to train Amazon AI models.
robots.txt is a request, not a lock: it only works for crawlers that choose to follow it, and it does not remove content already collected.
robots.txt is not for hiding or de-indexing pages
Anyone can read your robots.txt, so listing secret paths advertises them — protect private areas with a login. A URL blocked by robots.txt can still appear in Google results (without a description) if other pages link to it, because Google never fetches the page to see a noindex. To keep a page out of results, allow crawling and add <meta name="robots" content="noindex"> — the meta tag generator writes it for you.
Limitations
- The tester follows Google’s rules. Other crawlers mostly follow the same standard (RFC 9309) but can differ in details such as Crawl-delay or misspelt field names.
- The tool cannot download a live robots.txt from another site — browsers block reading other websites (CORS). Open the file in a new tab and paste it in.
- It cannot tell whether a visitor really is the crawler it claims to be; operators publish IP ranges or reverse-DNS checks for that.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
What is the difference between "Disallow:" and "Disallow: /"?
An empty Disallow: blocks nothing — the crawler may fetch everything. Disallow: / blocks the whole site, because every URL path starts with /.
Does Google support Crawl-delay?
No. Google ignores Crawl-delay and adjusts its crawl rate automatically based on how your server responds. Amazon’s crawlers ignore it too; some others, such as Anthropic’s ClaudeBot, honour it.
Will blocking GPTBot remove my site from ChatGPT search?
No. OpenAI says its robots.txt settings are independent: GPTBot covers training, while OAI-SearchBot decides whether pages can appear in ChatGPT search results. The same split exists at Anthropic (ClaudeBot vs Claude-SearchBot) and Amazon (Amazonbot vs Amzn-SearchBot).
Do I need a Sitemap line?
It is optional but useful: any crawler that reads robots.txt can discover your sitemap from it. Use the full URL, e.g. Sitemap: https://example.com/sitemap.xml; you can list several. Create one with the sitemap generator.
Does one robots.txt cover my subdomains?
No. Each protocol, host and port needs its own file: https://example.com/robots.txt does not apply to https://shop.example.com/ or to http://example.com/.
How quickly do changes take effect?
Google generally caches robots.txt for up to 24 hours, so allow about a day. Meta and Amazon also say updates can take around a day to be picked up.