Emergency server help: get in touch

Robots.txt Checker

Fetch and validate robots.txt, test any path for Googlebot, Bingbot and other search engines, and see which AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended and more) are allowed or blocked.

Status
Live
Last updated
October 3, 2026

Enter a domain to fetch its robots.txt and read it the way crawlers do. The checker validates the syntax, tests a path for Googlebot, Bingbot and other search engines, shows which AI crawlers are allowed or blocked and checks that every Sitemap line works.

How crawlers pick their rules

A crawler looks for the group whose User-agent line names it. Only when no group does, it follows the User-agent: * group; it never combines the two. Inside the group the longest matching Allow or Disallow rule wins, and Allow wins a tie (RFC 9309).

That is why adding a group for GPTBot with one Disallow line also removes every * rule for GPTBot. Repeat the rules you still want inside the named group.

AI crawlers: training versus answering

Training crawlers such as GPTBot, ClaudeBot, CCBot and Google-Extended collect text for AI models. Search and user agents such as OAI-SearchBot, Claude-SearchBot, PerplexityBot and ChatGPT-User fetch pages to answer a question and usually link to the source.

Many sites allow the answering agents and decide separately about training. Blocking Google-Extended does not affect Google Search, and blocking Applebot-Extended does not affect Siri or Spotlight.

Common robots.txt mistakes

Disallow: / left over from a staging site keeps the whole site out of search. A robots.txt that returns a server error makes Google pause crawling. Blocking CSS and JavaScript stops Google rendering your pages, and files over 500 KiB are cut off. Lines with typos are ignored silently; the checker lists them.

Robots.txt Checker at a glance

Robots.txt Checker summary card: Enter a domain to fetch its robots.txt and read it the way crawlers do.
In short: Enter a domain to fetch its robots.txt and read it the way crawlers do.
Robots.txt Checker sections: How crawlers pick their rules, AI crawlers: training versus answering and Common robots.txt mistakes
Covers: How crawlers pick their rules, AI crawlers: training versus answering and Common robots.txt mistakes.
Robots.txt Checker questions answered: Does robots.txt stop AI companies from using my content? Can robots.txt remove a page from Google?
Answers: Does robots.txt stop AI companies from using my content? Can robots.txt remove a page from Google?

Official documentation: RFC 9309 (Robots Exclusion Protocol), Google: robots.txt introduction.

Related guides: cPanel Temporary Domain: Free *.cpanel.site Setup · LiteSpeed Cache WordPress cPanel: Best Settings.

Frequently asked questions

Does robots.txt stop AI companies from using my content?

It stops crawlers that follow robots.txt, which includes the published crawlers of the large AI companies. It is a request, not access control, so crawlers that ignore it are not blocked.

Can robots.txt remove a page from Google?

No. A blocked page can still be indexed from links, without a description. Use a noindex meta tag or header on a page that crawlers are allowed to fetch.

What does the path field do?

It tests one URL path, such as /wp-admin/ or /shop/cart, against the rules so you can see exactly which rule applies to each crawler.

Free website test

Is your website set up right?

Check SSL, security headers, redirects, robots.txt, sitemap, llms.txt and security.txt in one test. It takes about 30 seconds.