My Tool Studio
SEO Tools·3 min read

How to Test Robots.txt Rules Against Real URLs

One wrong line in robots.txt can hide important pages from search. The Robots.txt Tester and Validator checks whether a URL is allowed or blocked for the crawler you pick and shows the exact rule that decided it. Load a site's live file or paste a draft, and test up to 100 URLs at once.

Robots.txt Tester and Validat…mytoolstudio.com › tools<title>…</title><meta name="description">SEO ready

How to test a robots.txt file

Load, list, read.

Three steps:

  • Choose Fetch from a site and press Load, or choose Paste robots.txt to test a draft. A fetched file appears in an editable box.
  • Enter one full URL or path per line, up to 100. Tick crawlers such as Googlebot, Bingbot or GPTBot, or type another user-agent token.
  • Read each verdict: Allowed or Blocked, the matching rule with its line number, and the user-agent group that applied. Select a result to highlight its rule in the line-numbered file view, and use Download CSV to share the table.

Finding the line behind a blocked page

The most common reason people arrive.

When Search Console reports a page as blocked by robots.txt and nobody can see why, paste that URL, tick Googlebot and the responsible line shows up in the results. Usually a rule matches more than expected. Disallow: /p also blocks /products and /pricing, because rules match from the start of the path.

Another common cause is a Googlebot group. A crawler obeys only the group that names it most specifically, so a User-agent: Googlebot group makes Google ignore everything under User-agent: *. Paths are case sensitive too: Disallow: /Admin/ does not block /admin/.

How the matching works

Google's rules, applied.

The tester follows Google's documented matching. The longest matching rule wins, and Allow wins a tie. The * wildcard matches any run of characters and $ marks the end of the URL, so Disallow: /*.pdf$ blocks /guide.pdf but not /guide.pdf?download=1.

It is not Google's own code, so treat the URL Inspection tool in Search Console as the final word.

Validating the file

Syntax problems, listed.

The same screen works as a validator. It lists rules placed before any User-agent line, missing colons, unknown or misspelled directives, and files over 500 KiB. When you load a live site, it also shows the HTTP status, any redirect and the content type of the file. A 404 means no rules, while a 5xx is treated by Google as a temporary block on the whole site. Check sitemap URLs tests each Sitemap line for its status and redirect target.

Are your CSS and scripts blocked?

The resources a page needs.

Google renders pages much like a browser, so blocked CSS or JavaScript can stop it seeing the page the way visitors do. Under Check a page's resources, enter a page URL and press Check resources. The tool fetches the page, lists every script, stylesheet, image and iframe it loads, and tests each one for the first crawler you ticked. Files on another host, such as a CDN, are tested against that host's own robots.txt.

Fixing and publishing

Test the fix before it goes live.

Edit the loaded file in the box and the results update as you type, so you can try a fix without touching the live site. Google generally caches robots.txt for up to 24 hours, and the robots.txt report in Search Console lets you request a recrawl.

Google retired its old robots.txt Tester, and the report that replaced it does not test URLs, so this page fills that gap. If you do not have a file yet, the Robots.txt Generator builds one and can send the draft straight here. Remember that Disallow stops crawling, not indexing. A noindex tag is the way to keep a page out of results.

Try it now

Open Robots.txt Tester and Validator

The tool is one click away. No sign up, no upload, no payment.

Open Robots.txt Tester and Validator