Robots.txt Generator
Build a valid robots.txt from a small form, with the right Allow and Disallow groups, a sitemap line and an optional block for AI crawlers, ready to copy.
Settings
Your robots.txt
A robots.txt file sits at the root of your site and tells crawlers which parts they may request. One stray slash can hide your whole site from Google, so the syntax is worth getting exactly right. This free generator builds a valid robots.txt from a small form: choose a preset, list the paths you want to keep out of crawling, point to your sitemap and, if you like, add dedicated groups that block the AI crawlers. The output updates live and is ready to copy into a file named robots.txt. It is built for SEOs, developers and site owners who want a correct file without memorising the format.
How to use it
Pick a preset
Start from Allow all crawlers, Block entire site or Custom rules. The first two write the whole file for you; Custom opens the user-agent and disallow fields so you can shape it.
Add your rules
On Custom, set the user-agent (leave it as * for every crawler) and list the paths to disallow, one per line, like /admin/ or /cart. Add your sitemap URL and tick the box to block AI crawlers if you want to.
Copy it to your site
Copy the generated text, save it as a plain file named robots.txt and upload it to the root of your domain so it is reachable at /robots.txt. Crawlers read it from there.
Why it's useful
Valid syntax, every time
The generator writes the groups in the order crawlers expect: a user-agent line, its Allow and Disallow lines, a blank line between groups and the Sitemap line at the end. You avoid the small mistakes, a missing colon or a stray space, that quietly break a hand-typed file.
Block AI crawlers in one click
A single checkbox appends dedicated Disallow groups for the known AI crawlers, from GPTBot and ClaudeBot to CCBot, Google-Extended and PerplexityBot, so you can keep your content out of AI training and retrieval without looking up each agent name.
Nothing leaves your browser
The file is built entirely client-side. Nothing you type is uploaded or stored, so you can draft rules for an unreleased site or a client project without signing in.
How robots.txt controls crawling, and where it does not
What robots.txt is and how crawlers read it
Robots.txt is a plain text file at the root of your domain, reachable at /robots.txt, that follows the Robots Exclusion Protocol. A crawler fetches it before it requests your pages and obeys the rules that apply to its user-agent. The file is organised in groups: a User-agent line names the crawler, and the Disallow and Allow lines under it say which paths that crawler may or may not request. A blank line separates one group from the next, and a Sitemap line, which is not tied to any group, usually sits at the end.
The matching is simpler than it looks. Disallow: /admin/ blocks every URL whose path starts with /admin/. An empty Disallow line means nothing is blocked, which is how you allow everything, while Disallow: / blocks the entire site. A User-agent of * is the default group that applies to any crawler without its own, and most sites need only that one group plus a sitemap line.
Blocking crawling is not blocking indexing
The most common misunderstanding is that Disallow keeps a page out of Google. It does not. Disallow stops a crawler from fetching the page, but a URL that is linked from elsewhere can still be indexed without its content, showing up as a bare link with no description. If your goal is to keep a page out of the results entirely, let it be crawled and add a noindex meta tag, or protect it behind authentication. Robots.txt controls crawling; it does not control indexing.
This also means robots.txt is the wrong tool for private data. Because the file is public, anyone can read your Disallow list, so listing a secret folder there effectively advertises it. Use it to steer crawl budget away from low-value or duplicate URLs, like internal search results, faceted filters or a cart, and use authentication or noindex for anything that must genuinely stay hidden.
Do not block your own assets
A classic mistake is disallowing the folders that hold your CSS and JavaScript. Google renders pages to understand them, and if it cannot fetch those files the page may be judged on a broken rendering. Keep your assets crawlable and reserve Disallow for pages, not the resources that make them work.
AI crawlers and the limits of the file
A newer use of robots.txt is telling AI crawlers to stay away. Agents like GPTBot, ClaudeBot, CCBot, Google-Extended and PerplexityBot each read their own user-agent group, so blocking them means adding a Disallow group per agent, which the AI-crawler checkbox here does for you. Remember that robots.txt is a request, not a wall: well-behaved crawlers obey it, but it is enforced by convention, not by the server, so anything that must be truly inaccessible needs real access control. If instead your aim is to be found and cited across every language you publish in, that is where Otorank picks up, researching, writing, publishing and tracking content built to be found across 150+ languages once you move from one file to running the whole channel.
They use Otorank
Questions
What robots.txt can and cannot do, where the file goes, and how blocking AI crawlers works.
At the root of your domain, so it is reachable at https://yourdomain.com/robots.txt. Crawlers only look there; a robots.txt in a subfolder is ignored.
No. Disallow stops the page being crawled, but a URL linked from elsewhere can still be indexed without its content. To keep a page out of results, allow crawling and use a noindex meta tag, or require authentication.
An empty Disallow line blocks nothing, which is how you allow everything for that user-agent. Disallow: / with a slash blocks the entire site, so the single character makes all the difference.
Add a group per AI user-agent with Disallow: /. Tick Block AI crawlers here and the generator appends those groups for GPTBot, ClaudeBot, CCBot, Google-Extended, PerplexityBot and others automatically.
No. The file is public and only requests compliance, so it cannot protect sensitive content. Use authentication or a noindex tag for anything that must stay hidden, and reserve robots.txt for steering crawl budget.