Skip to content
↑↓ to move ↵ to open Esc to close Browse all tools

Robots.txt Generator

Build a robots.txt file with rules per crawler, your sitemap and an optional block for AI crawlers, then test which URLs are allowed.

Robots.txt generator

Default for all crawlers
Applies to User-agent: *, the group for every crawler that has no group of its own.

Rules

Disallow blocks a path and everything below it. Use * for any characters and $ for the end of the URL.

Common paths to block

Added to the * group. The WordPress option keeps /wp-admin/admin-ajax.php allowed, because many themes and plugins need it.

AI crawlers

Google ignores Crawl-delay. Bing and some other crawlers use it.
Your robots.txtUpload it to the root of your site
 

Test a URL

Checks the file above the way Google reads it: the most specific user-agent group, the longest matching rule, and Allow when two rules are equally long.

The file is built in your browser. Nothing you enter is sent to a server.

What robots.txt does, and what it does not

A robots.txt file is a plain text file at the root of your site that tells crawlers which URLs they may fetch. Search engines like Google and Bing read it before they crawl, and so do many other bots. You use it to keep crawlers away from pages that waste their time, such as internal search results, endless filter combinations, cart and checkout pages or admin areas, and to point them to your sitemap.

What robots.txt does not do is hide pages. The file is public, anyone can read it, and a disallowed URL can still appear in search results if other pages link to it, just without a description. To keep a page out of Google, let it be crawled and add a noindex meta tag or an X-Robots-Tag header. If Google cannot crawl the page, it never sees the noindex. Anything truly private belongs behind a login.

How crawlers read the rules

  • Groups. Each group starts with one or more User-agent lines followed by Allow and Disallow rules.
  • One group per crawler. A crawler follows only the most specific group that names it, so Googlebot ignores the * rules once it has a group of its own. Repeat shared rules in every group that needs them.
  • Longest match wins. When several rules match a URL, Google uses the one with the longest path. If an Allow and a Disallow rule are equally long, Allow wins.
  • Wildcards. * stands for any characters, and $ marks the end of the URL. Rules are case sensitive.
  • Sitemaps. Sitemap: lines take full URLs and apply to all crawlers, no matter where they appear in the file.

The tester above follows these rules, so you can check a tricky URL before you upload the file.

Robots.txt examples

Block a folder but allow one page inside it, and list the sitemap:

User-agent: *
Disallow: /private/
Allow: /private/press-kit.html

Sitemap: https://example.com/sitemap.xml

Block all PDF files for every crawler:

User-agent: *
Disallow: /*.pdf$

Block a staging site completely. Use a password as well, because a blocked site can still show up as bare URLs:

User-agent: *
Disallow: /

Common AI crawler user-agents

Many AI companies publish the name their crawler uses, so you can block them in robots.txt. These are the ones the generator lists:

User-agentOperated by
GPTBotOpenAI
ChatGPT-UserOpenAI
OAI-SearchBotOpenAI
ClaudeBotAnthropic
anthropic-aiAnthropic
Google-ExtendedGoogle
CCBotCommon Crawl
PerplexityBotPerplexity
BytespiderByteDance
Applebot-ExtendedApple

Some of these collect training data, others fetch pages for AI search or when a user asks an assistant about a link. Check each company's documentation for what its bot does, and review the list from time to time, since new crawlers appear and names change. robots.txt only works with bots that choose to respect it. To block a bot that ignores it, you need rules on your server, CDN or firewall.

How to use the robots.txt generator

  1. 1Choose whether crawlers may crawl everything by default, then add disallow and allow paths for all crawlers or add a group for a specific bot such as Googlebot or Bingbot.
  2. 2Tick common paths to block, block AI crawlers if you want to, and add your sitemap URL.
  3. 3Test a few important URLs, then copy or download the file and upload it as robots.txt to the root of your site. Check it in Google Search Console afterwards.

Frequently asked questions

Where do I put the robots.txt file?

In the root folder of your site, so it loads at https://example.com/robots.txt. The file name must be lowercase and it must be plain text. Crawlers do not look for it in subfolders, and every subdomain, like shop.example.com, needs its own file.

Does robots.txt keep a page out of Google?

Not reliably. Disallow stops crawling, not indexing. If other sites link to a blocked page, Google can still list its URL without a description. To keep a page out of search results, allow crawling and add a noindex meta tag or an X-Robots-Tag header. For private content, use a password.

Should I block AI crawlers?

That is a business decision. Blocking crawlers that collect training data, like GPTBot, ClaudeBot, CCBot, Google-Extended or Applebot-Extended, does not affect your rankings in Google Search. Blocking bots used for AI search and assistants, like OAI-SearchBot, ChatGPT-User or PerplexityBot, can keep your pages out of answers that would link to you. The list of bots changes often, and robots.txt is a request that not every crawler follows.

What do * and $ mean in a rule?

The asterisk matches any run of characters, and a dollar sign at the end means the URL must end there. Disallow: /*.pdf$ blocks every URL that ends in .pdf, while /*.pdf without the dollar sign also blocks /file.pdf?download=1. Paths are case sensitive, so /Private/ and /private/ are different rules.

How fast does Google pick up a new robots.txt?

Google says it generally caches robots.txt for up to 24 hours, so changes can take a day to apply. In Google Search Console, the robots.txt report shows which version Google fetched last and lets you ask for a recrawl after an urgent fix.