Robots.txt Tester
Paste your robots.txt and test whether any URL is crawlable for a given user-agent, with the exact rule that decides. Free, client-side, no account.
Your robots.txt
A robots.txt file tells search-engine crawlers which paths they are allowed to fetch, and a single misplaced Disallow can quietly keep an important page from being crawled. This free robots.txt tester parses your file the way Googlebot does, picks the matching user-agent group, and tells you whether each URL is allowed or blocked, naming the exact rule that decides. It is built for SEOs and site owners who want to confirm what a crawler can reach before a change ships, so the right pages stay visible and get found.
How to use it
Paste your robots.txt
Drop in the contents of your robots.txt file. Any sitemap lines are detected and listed, and the groups and rules are parsed live as you type.
Choose the user-agent
Enter the crawler you want to test, such as Googlebot or *. The tester selects the most specific group that names it, exactly as a real crawler would.
Test your URLs
Add one URL or path per line. Each gets an Allowed or Blocked verdict alongside the precise Allow or Disallow rule that produced it.
Why it's useful
See the exact rule that decides
Instead of guessing, you see which Allow or Disallow line won and why. The tester applies Google's longest-match precedence, with Allow beating Disallow on a tie, so the verdict matches how Googlebot reads the same file.
Matches real crawler behaviour
A crawler obeys only the single most specific user-agent group that names it, never a merge of several. This tool picks the group the same way, so a Googlebot-specific block is not hidden behind the catch-all group you were reading.
Nothing leaves your browser
The whole check runs client-side. Your robots.txt and test URLs are never uploaded or stored, so you can safely test staging files and client sites without signing in.
How robots.txt is actually read by a crawler
Crawling is not indexing
The most common robots.txt mistake comes from a single misunderstanding: robots.txt controls crawling, not indexing. It tells compliant crawlers which paths they may fetch, nothing more. A page you block in robots.txt can still appear in search results if other pages link to it, because Google can index the URL from those links without ever fetching the page. It just shows the result without a description, since it was never allowed to read the content.
That means blocking a page in robots.txt is the wrong tool for keeping it out of the index, the opposite of what most people expect. To actually remove a page from search, let the crawler fetch it and add a noindex meta tag or X-Robots-Tag header. If the URL is disallowed, Google never sees the noindex, so the page lingers. Use robots.txt to save crawl budget on low-value paths, and use noindex to control what appears in the index.
How a crawler picks the group and the rule
Rules are grouped by user-agent, and a crawler obeys only the most specific group that names it. If there is a dedicated Googlebot group, Googlebot ignores the * group entirely, it does not combine the two. This tester resolves the group the way a crawler does: the longest user-agent token that matches your crawler wins, falling back to * when none name it. That alone explains many surprises, where a rule in the catch-all group never applies to the bot you care about.
Longest match wins, Allow breaks ties
Within the chosen group, when several rules match a URL the most specific one wins, measured by the length of the path it matches, and Allow beats Disallow when the lengths tie. So Disallow: /cart together with Allow: /cart/shared lets /cart/shared through while still blocking /cart. The tester shows you exactly which line decided each URL, so you can see why something is allowed or blocked rather than reading the file by eye.
Wildcards, the public file, and testing what ships
Two wildcards are supported in path patterns: * matches any run of characters, and $ anchors the end of the URL. So Disallow: /*.pdf$ blocks every URL ending in .pdf, while Disallow: /search? blocks anything under your search endpoint. Keep patterns as simple as the job needs, because a long chain of wildcards is hard to reason about and easy to get subtly wrong.
Remember that robots.txt is a public file served at /robots.txt, readable by anyone, so never list secret or sensitive paths in it, that only advertises where they are. Finally, test the file you actually ship. A robots.txt that looks right in your repo can differ from the one your server returns in production, where a CDN rule or a stray redirect has changed it. Paste the live, rendered file into this tester so you validate what crawlers really get.
They use Otorank
Questions
How robots.txt precedence works, why a blocked page can still rank, and what this tester does and does not check.
No. It only controls crawling. A disallowed URL can still appear in results if other pages link to it, shown without a description. To keep a page out of the index, allow it to be crawled and add a noindex tag or header instead.
Within the matching user-agent group, the rule whose path matches the most characters wins, and Allow beats Disallow on a tie. This tester applies that same longest-match, Allow-wins precedence and names the deciding rule for each URL.
A crawler follows only the single most specific group that names it. If a Googlebot group exists, Googlebot uses it exclusively and never merges in the * group, so any catch-all rules simply do not apply to it.
Yes. * matches any sequence of characters and $ anchors the end of the URL, so Disallow: /*.pdf$ blocks every PDF. The tester evaluates both when deciding whether a path matches a rule.
No. Everything is parsed in your browser; nothing is transmitted or stored, so you can safely test staging files and client sites without signing in.