Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

Indexability Checker

Can Google index this URL? One check, the exact reason.

SEO Uses live data Free, no sign-up

Check a URL

The exact address you want in search results. MySmartCoPilot’s server fetches it, follows redirects (up to 10, like Googlebot) and reads robots.txt for every host on the way.

Send the request as

Next steps

About the Indexability Checker

Enter a URL and the checker works through everything that decides whether a search engine may index it, in the order a crawler meets it: the robots.txt rules for its host (for Googlebot and Bingbot), the HTTP status and every redirect, the X-Robots-Tag header, robots meta tags, the canonical URL from the HTML and the Link header, and the page’s hreflang alternates.

It ends with one verdict per crawler and the exact reason — the robots.txt line that blocks it, the header or tag that says noindex, the status code, or the canonical that points somewhere else — so you know what to fix. Rules follow Google Search Central’s documentation and RFC 9309, the robots.txt standard; the robots.txt matching reuses MySmartCoPilot’s robots.txt tester.

How to use it

  1. Paste the exact URL you want in search results, including https:// and www if your site uses them.
  2. Leave Send the request as on MySmartCoPilotBot. Choose Mobile browser if the site blocks unknown bots and you want to see what a phone browser gets.
  3. Press Check indexability. The two verdict cards answer the question for Googlebot and Bingbot; the checks table shows each signal side by side.
  4. Read What this means for the details and fixes, and open the robots.txt, redirect and header sections to see exactly what the server sent.
  5. Copy or download the report to share it, then check again after you fix the problem.

Examples

A page blocked by robots.txt
Input
https://www.example.com/private/report
Result
Googlebot: Blocked by robots.txt — Blocked by “Disallow: /private/” (line 4) in the “User-agent: *” group.

Googlebot won’t fetch the page, so it never sees a noindex on it. The bare URL can still be listed (without a description) if other sites link to it.

noindex sent only to Google
Input
X-Robots-Tag: googlebot: noindex
Result
Googlebot: noindex in the X-Robots-Tag header
Bingbot: Indexable

A user agent before the rules limits them to that crawler. Without one, the rules apply to every crawler.

A canonical that points to the www version
Input
https://example.com/chai/ with <link rel="canonical" href="https://www.example.com/chai/">
Result
Canonical points to another URL — differs by www vs non-www.

Search engines will probably index the www URL instead. If that is intended, also redirect the non-www version to it.

Common uses

  • Finding out why a page is missing from Google: robots.txt, noindex, a 4xx/5xx status or a canonical to another page.
  • Checking a new section, staging copy or migrated site before launch — the classic leftover noindex or Disallow: /.
  • Making sure PDFs and other files you want hidden send X-Robots-Tag: noindex, and that the ones you want found don’t.
  • Spotting rules that treat Googlebot and Bingbot differently, such as a User-agent: Googlebot group.

What decides whether a URL can be indexed

  • robots.txt controls crawling. A blocked URL is not fetched at all, so the crawler never sees its status, redirects or noindex. Google may still index the bare URL if other pages link to it.
  • The HTTP status. Google doesn’t index URLs that return 4xx (except 429); 429 and 5xx are server errors that slow crawling and eventually drop the URL; a 2xx doesn’t guarantee indexing. Googlebot follows up to 10 redirect hops and processes the target’s content, not the redirecting URL’s.
  • noindex in the X-Robots-Tag header or a robots meta tag keeps the page out of results. When rules conflict, the more restrictive one applies, and Google respects robots meta tags in the <body> too.
  • The canonical. rel="canonical" (in the <head> only) or a Link: <…>; rel="canonical" header is a strong signal; if it names another URL, that URL is usually indexed instead.
  • Redirects inside the page. An instant meta refresh or Refresh: 0; url=… header counts as a permanent redirect and a delayed one as temporary; a JavaScript redirect only works when Google manages to render the page. The checker reports literal location redirects it finds in inline scripts, without running them.

Sources: Google Search Central — robots.txt specification, Robots meta tag and X-Robots-Tag, HTTP status codes, canonical URLs and redirects.

How robots.txt answers are treated

  • 200 — the rules apply. Google reads at most 500 KiB of the file.
  • Redirects — Google follows at least five hops, then treats robots.txt as missing.
  • 4xx (except 429) — treated as “no robots.txt”: nothing is blocked. Don’t use 401 or 403 to limit crawling.
  • 5xx, 429, timeouts and connection errors — Google stops crawling the whole site for the first 12 hours, then uses the last good copy of robots.txt for up to 30 days (or assumes no restrictions if it has none).
  • Bingbot is evaluated with RFC 9309: 4xx means crawlers may crawl everything, and an unreachable robots.txt (5xx) means they must assume everything is disallowed.

Rules are matched like Google’s parser: the crawler’s own group (or *), the longest matching rule wins, and Allow wins a tie.

What the checker sends

Requests come from Cloudflare’s network with the user agent MySmartCoPilotBot/1.0 (or a mobile Chrome string if you choose it) and no cookies. Search-engine crawlers are never impersonated. Before each new host is contacted, its DNS records are looked up over HTTPS and every address must be public; private, loopback, link-local and cloud-metadata addresses are never fetched, and every redirect is checked again. Each request times out after 8 seconds. At most 1 MiB of the page and 500 KiB of each robots.txt are read, and robots.txt is fetched for at most three hosts per run.

Limitations

  • JavaScript is not run. Google renders pages, so a canonical or robots tag added by JavaScript can change the result — though a noindex in the original HTML may stop Google rendering the page at all.
  • Bing doesn’t publish rules as detailed as Google’s, so Bingbot is evaluated with RFC 9309 and the robots and bingbot meta names, and Google-only rules such as unavailable_after are not applied to it. Rules for other crawlers are listed but not evaluated.
  • Sites can answer MySmartCoPilot’s server differently from Googlebot (bot protection, country or device redirects, cloaking). Search Console’s URL Inspection shows what Google itself fetched.
  • A canonical that points elsewhere is checked against robots.txt but not fetched, and hreflang alternates are not fetched either — so a canonical or alternate that redirects or returns an error is only caught when you check that URL too.
  • “Indexable” means nothing blocks indexing — not that the page is indexed. Search engines still choose what to index, and soft 404s (error pages that return 200) are only flagged by simple signs such as an empty page or a “not found” title.
  • When a server sends several X-Robots-Tag headers they arrive joined by commas, so rules after a user agent are read as addressed to that user agent.
  • Each visitor can run a limited number of checks per hour (and the whole site a limited number per day) to stay within free hosting limits.

Privacy

The URL you enter is sent to MySmartCoPilot’s server, which requests it and the robots.txt files from the website. MySmartCoPilot does not store the URL or the result. For abuse prevention, the server log records only the host name, status code and time of each request — never the full URL or your IP address.

Frequently asked questions

Why does a page blocked by robots.txt still show up in Google?

robots.txt stops crawling, not indexing. If other pages link to the URL, Google can index the bare address without fetching it, usually with no description. To keep a page out of results, allow crawling and use noindex — the crawler has to fetch the page to see it.

What is the difference between the robots meta tag and X-Robots-Tag?

They accept the same rules. The meta tag goes in a page’s HTML (<meta name="robots" content="noindex">); the X-Robots-Tag HTTP header also works for PDFs, images and other files. Either can be aimed at one crawler — <meta name="googlebot" …> or X-Robots-Tag: googlebot: noindex.

My robots.txt returns an error. Does that matter?

A 404 is fine: crawlers treat it as “no rules”. A 5xx, 429 or timeout is serious: Google stops crawling the whole site for the first 12 hours and then relies on an old copy. Make robots.txt return 200 (or 404 if you have no rules).

Does content="noindex nofollow" without a comma work?

Google documents robots rules as a comma-separated list (noindex, nofollow) and says nothing about spaces or semicolons, so the checker reports such a value as unclear instead of guessing. Add the comma if you want the page kept out of results, or remove noindex if you want it indexed.

Is a canonical tag a directive?

No. Google describes rel="canonical" and redirects as strong signals and sitemaps as a weak one; it chooses the canonical itself. A canonical to another URL usually means that URL is indexed instead, so make sure every page’s canonical names the URL you want in results.

Why is the result for Bingbot different from Googlebot?

Most often robots.txt has a group just for one of them (User-agent: Googlebot or User-agent: Bingbot), or a meta tag or header names one crawler. robots.txt errors also differ: Google pauses crawling for 12 hours, while RFC 9309 tells crawlers to treat everything as disallowed.

Does it check the whole website?

No — one URL per run, plus the robots.txt of each host it passes through. To check many pages, use the sitemap validator, which can test the status of a sample of URLs from your sitemap.

Why was my URL refused?

For security the checker only fetches public websites on ports 80 and 443. Private or local addresses (such as 192.168.x.x or 127.0.0.1), internal host names and MySmartCoPilot’s own pages are refused, and every redirect target is checked again before it is followed.

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.