Free SEO tools

Robots.txt Generator

Build a valid robots.txt from a small form, with the right Allow and Disallow groups, a sitemap line and an optional block for AI crawlers, ready to copy.

Settings

Your robots.txt

 

A robots.txt file sits at the root of your site and tells crawlers which parts they may request. One stray slash can hide your whole site from Google, so the syntax is worth getting exactly right. This free generator builds a valid robots.txt from a small form: choose a preset, list the paths you want to keep out of crawling, point to your sitemap and, if you like, add dedicated groups that block the AI crawlers. The output updates live and is ready to copy into a file named robots.txt. It is built for SEOs, developers and site owners who want a correct file without memorising the format.

How to use it

1

Pick a preset

Start from Allow all crawlers, Block entire site or Custom rules. The first two write the whole file for you; Custom opens the user-agent and disallow fields so you can shape it.

2

Add your rules

On Custom, set the user-agent (leave it as * for every crawler) and list the paths to disallow, one per line, like /admin/ or /cart. Add your sitemap URL and tick the box to block AI crawlers if you want to.

3

Copy it to your site

Copy the generated text, save it as a plain file named robots.txt and upload it to the root of your domain so it is reachable at /robots.txt. Crawlers read it from there.

Why it's useful

Valid syntax, every time

The generator writes the groups in the order crawlers expect: a user-agent line, its Allow and Disallow lines, a blank line between groups and the Sitemap line at the end. You avoid the small mistakes, a missing colon or a stray space, that quietly break a hand-typed file.

Block AI crawlers in one click

A single checkbox appends dedicated Disallow groups for the known AI crawlers, from GPTBot and ClaudeBot to CCBot, Google-Extended and PerplexityBot, so you can keep your content out of AI training and retrieval without looking up each agent name.

Nothing leaves your browser

The file is built entirely client-side. Nothing you type is uploaded or stored, so you can draft rules for an unreleased site or a client project without signing in.

How robots.txt controls crawling, and where it does not

What robots.txt is and how crawlers read it

Robots.txt is a plain text file at the root of your domain, reachable at /robots.txt, that follows the Robots Exclusion Protocol. A crawler fetches it before it requests your pages and obeys the rules that apply to its user-agent. The file is organised in groups: a User-agent line names the crawler, and the Disallow and Allow lines under it say which paths that crawler may or may not request. A blank line separates one group from the next, and a Sitemap line, which is not tied to any group, usually sits at the end.

The matching is simpler than it looks. Disallow: /admin/ blocks every URL whose path starts with /admin/. An empty Disallow line means nothing is blocked, which is how you allow everything, while Disallow: / blocks the entire site. A User-agent of * is the default group that applies to any crawler without its own, and most sites need only that one group plus a sitemap line.

Blocking crawling is not blocking indexing

The most common misunderstanding is that Disallow keeps a page out of Google. It does not. Disallow stops a crawler from fetching the page, but a URL that is linked from elsewhere can still be indexed without its content, showing up as a bare link with no description. If your goal is to keep a page out of the results entirely, let it be crawled and add a noindex meta tag, or protect it behind authentication. Robots.txt controls crawling; it does not control indexing.

This also means robots.txt is the wrong tool for private data. Because the file is public, anyone can read your Disallow list, so listing a secret folder there effectively advertises it. Use it to steer crawl budget away from low-value or duplicate URLs, like internal search results, faceted filters or a cart, and use authentication or noindex for anything that must genuinely stay hidden.

Do not block your own assets

A classic mistake is disallowing the folders that hold your CSS and JavaScript. Google renders pages to understand them, and if it cannot fetch those files the page may be judged on a broken rendering. Keep your assets crawlable and reserve Disallow for pages, not the resources that make them work.

AI crawlers and the limits of the file

A newer use of robots.txt is telling AI crawlers to stay away. Agents like GPTBot, ClaudeBot, CCBot, Google-Extended and PerplexityBot each read their own user-agent group, so blocking them means adding a Disallow group per agent, which the AI-crawler checkbox here does for you. Remember that robots.txt is a request, not a wall: well-behaved crawlers obey it, but it is enforced by convention, not by the server, so anything that must be truly inaccessible needs real access control. If instead your aim is to be found and cited across every language you publish in, that is where Otorank picks up, researching, writing, publishing and tracking content built to be found across 150+ languages once you move from one file to running the whole channel.

They use Otorank

FAQ

Questions

What robots.txt can and cannot do, where the file goes, and how blocking AI crawlers works.

At the root of your domain, so it is reachable at https://yourdomain.com/robots.txt. Crawlers only look there; a robots.txt in a subfolder is ignored.

No. Disallow stops the page being crawled, but a URL linked from elsewhere can still be indexed without its content. To keep a page out of results, allow crawling and use a noindex meta tag, or require authentication.

An empty Disallow line blocks nothing, which is how you allow everything for that user-agent. Disallow: / with a slash blocks the entire site, so the single character makes all the difference.

Add a group per AI user-agent with Disallow: /. Tick Block AI crawlers here and the generator appends those groups for GPTBot, ClaudeBot, CCBot, Google-Extended, PerplexityBot and others automatically.

No. The file is public and only requests compliance, so it cannot protect sensitive content. Use authentication or a noindex tag for anything that must stay hidden, and reserve robots.txt for steering crawl budget.

Related free tools

FREE TRIAL

Ready to become
visible on Google?

Your SEO visibility, on autopilot

Start by adding your website

We'll scan your site and show opportunities to get visible on Google.

Get set up in 60 seconds No credit card required Cancel anytime