Robots Txt Generator
Builds a robots.txt file from user-agent groups, Allow and Disallow paths, Sitemap lines, and optional Crawl-delay values. Presets block AI training or AI search crawlers by their documented tokens, and paths that do not start with / or user-agents that are not RFC 9309 product tokens are flagged. A group with no rules is written as an empty Disallow, which permits everything.
This tool sends nothing over the network. Everything you enter is processed on your device and never reaches our servers.
Documentation
A robots.txt generator writes the plain-text file that tells crawlers which paths on a host they may fetch. The file lives at the root of the host, such as https://example.com/robots.txt, and the format has been an Internet standard since RFC 9309 was published in September 2022. Each protocol, host, and port needs its own file, so https://shop.example.com is governed by a different robots.txt than https://example.com. The file controls crawling, not indexing: a disallowed URL can still appear in search results when other pages link to it, and a noindex tag only works on a page the crawler is allowed to fetch.
Rules are written in groups. A group opens with one or more User-agent lines naming a crawler by its product token, then lists Allow and Disallow rules. A crawler picks the group whose token matches its own name, compared without regard to case, and merges every group that names it. Only a crawler with no group of its own falls back to the group for *, so naming GPTBot in a separate group means GPTBot ignores the rules written for everyone else. RFC 9309 tokens contain only letters, hyphens, and underscores, which is why a full browser-style user agent string never matches anything and is flagged here.
A rule path starts with / and matches as a prefix: Disallow: /private/ covers /private/report and everything else under that directory, but not /private on its own. Two characters are special. An asterisk stands for any run of characters, and a dollar sign anchors the end of the URL. When several rules match the same URL, the one with the longest path, counted in bytes, decides; if an Allow and a Disallow rule are the same length, the Allow wins. A group with no rules is written as an empty Disallow line, which permits everything, while Disallow: / blocks the whole host.
Two other lines appear in most files. Sitemap lines are read independently of the groups and must carry a full URL. Crawl-delay asks for a pause in seconds between requests; it is outside RFC 9309, Bingbot and Anthropic's crawlers read it, and Googlebot ignores it. Crawlers must read at least the first 500 kibibytes of the file, may keep a cached copy for up to 24 hours, treat a 4xx response as permission to crawl everything, and treat a 5xx response as a complete block. Settings holds a default Crawl-delay for groups without their own value and the option to write comment lines that label each group.
The AI presets list tokens from each operator's own documentation. GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, and Bytespider collect training data; OAI-SearchBot, Claude-SearchBot, PerplexityBot, and Meta-WebIndexer fetch pages for AI search answers. Google-Extended and Applebot-Extended are control tokens rather than crawlers: Google and Apple read them to decide whether content their search crawlers already fetched may train their models, so blocking them leaves pages in Google Search and Apple's search results. Fetchers that act on a single user's request, such as ChatGPT-User, Perplexity-User, and Meta-ExternalFetcher, are left out because their operators state that robots.txt may not apply to them. Every rule is a request that well-behaved crawlers honor; a crawler that ignores the file can only be stopped by blocking it at the server.
Take one group for every crawler with three rules: Disallow: /shop/, Allow: /shop/sale/, and Disallow: /*.pdf$. A request for /shop/sale/boots matches both shop rules, and Allow: /shop/sale/ wins because its path is 11 bytes long against 6 for /shop/. A request for /shop/cart matches only Disallow: /shop/ and is blocked, while /shop itself matches neither rule and stays open. The PDF rule blocks /docs/guide.pdf, but /docs/guide.pdf?v=2 is allowed, because the dollar sign requires the URL to end right after .pdf.
A correctly configured robots.txt file is a routine part of search engine optimization. It steers crawlers away from low-value paths, keeps crawl activity focused on public pages, and records which automated agents a site owner has asked to stay out. The scenarios below are the common reasons to write or revise one.
- New Site Launch: Define a baseline crawl policy before a domain goes live. Keep crawlers out of leftover /staging/ and /dev/ directories while allowing all production content, and add the sitemap URL so search engines find the full page list from the first visit.
- AI Crawler Management: Ask training crawlers such as GPTBot, ClaudeBot, and CCBot to stay out, and add the Google-Extended control token to keep content out of Google's Gemini models, while ordinary search crawlers keep full access. The two AI presets separate training from AI search, so a site can stay visible in AI answers while declining training use.
- E-commerce Crawl Budget: Large catalogs generate endless filtered and sorted URLs. Disallowing /search, /cart/, and /checkout/ keeps crawlers on product and category pages; because every rule is a prefix match, a trailing asterisk adds nothing.
- Content Management Systems: WordPress's default file disallows /wp-admin/ but allows /wp-admin/admin-ajax.php, a working example of the longer Allow path overriding the shorter Disallow for a single file that front-end features depend on.
- Multi-Region Sites: Blocking duplicate locale paths stops them from being crawled, but it also hides their hreflang and canonical tags from search engines, so annotations are usually the better fix and robots.txt the last resort for paths that should never be fetched.
- API and Development Environments: Restrict crawlers from API endpoints, sandbox environments, and internal tooling on public-facing subdomains, each of which needs its own robots.txt at its own root.
- Media and Publishing: Keep crawlers away from draft directories and print-only duplicates while headline and article pages stay crawlable. Paywalled content needs a server-side control, since robots.txt does not restrict human visitors.
- Periodic Audit: Review and regenerate the file when site structure changes, new directories are added, or crawl patterns shift after a redesign, and check that no old Disallow rule now covers pages that should rank.
- Staging Lockdown: A staging host can carry Disallow: / during development and switch to a permissive file at launch. Because robots.txt does not stop indexing of linked URLs, password protection is the safer guard for staging content.
Inputs, outputs, and what the Robots Txt Generator computes
What the Robots Txt Generator asks for and what it returns, as a plain list. Defaults, units, and ranges are the ones the form loads with.
Inputs
- Start from a preset · default: Custom (start blank)
- Include explanatory comments in output · default: on
- Default Crawl-delay (seconds, leave blank to omit) (text input)
Controls
Generate · Reset · Copy to Clipboard · Download robots.txt
Example
A request for /shop/sale/boots matches both shop rules, and Allow: /shop/sale/ wins because its path is 11 bytes long against 6 for /shop/.