Robots.txt Generator & Robots.txt Validator

FREE
SEO/Marketing

Create or audit your crawler rules with one free tool. Our online robots txt generator lets you configure User-agent, Allow, Disallow, Sitemap, Crawl-delay, AI-bot access, and directory restrictions. The built-in Robots.txt Validator checks a live file and summarizes sitemap detection, broad crawler blocks, crawlability, and issues that may need attention. Generate a server-ready file, copy the code, or download it as a .txt file. Already have robots.txt? Enter the live URL to review the accessible file and its detected directives before making changes.

Create and Check Robots.txt Without Writing Every Rule by Hand

A robots.txt file gives compliant crawlers instructions about which paths they may request. This tool helps you build those instructions and then analyze the live file from the same page. It is designed for site owners, developers, SEO teams, publishers, and marketers who need clear crawl controls without manually formatting every directive.

Use the generator when you are creating a new file or replacing an outdated one. Use the validator when you want a quick site-level view of the live robots.txt file, sitemap reference, broad crawler restrictions, crawlability, and detected recommendations.

How to Use the Online Robots Txt Generator

  1. Choose the crawler or User-agent. Apply a rule to all compliant crawlers with an asterisk, or select a specific crawler option available in the tool.
  2. Set Allow and Disallow paths. Define the pages or directories that crawlers may request or should avoid. Paths are case-sensitive, so enter them exactly as they appear in your URLs.
  3. Add your sitemap. Enter the complete sitemap URL, including the protocol and hostname for example, https://example.com/sitemap.xml.
  4. Configure optional controls. Add Crawl-delay for crawlers that support it, choose AI-bot crawling controls, and add directory restrictions needed for your site.
  5. Review the generated code. Check every path before publishing. A single sitewide Disallow rule can prevent compliant crawlers from accessing the whole site.
  6. Copy or download the file. Copy the output to your clipboard or download the generated robots.txt file.
  7. Upload it to the site root. Publish the file at https://yourdomain.com/robots.txt for the exact host it should control.
  8. Run the validator. Enter the live URL after deployment to confirm that the file is accessible and review the reported status cards and recommendations.

How to Use the Robots.txt Validator

The validator analyzes a live robots.txt URL rather than testing one individual page against a selected user agent. Enter your website or robots.txt address, select Check, and review the file-level results.

  • Sitemap Found: Shows whether the fetched robots.txt file contains a detectable Sitemap directive.
  • Admin Blocked: Highlights whether a broad rule appears to block crawler access across the site. Review any sitewide Disallow rule carefully before publishing.
  • Crawlability: Provides a high-level assessment of how open or restrictive the detected rule set appears.
  • ROBOTS.TXT tab: Displays the retrieved directives with line-level readability so you can inspect the live file.
  • SITEMAP tab: Lets you review the sitemap information detected during the analysis.
  • Issues & Recommendations: Summarizes positive checks and potential concerns, such as overall site crawlability and sitemap-directive presence.
Scope of the validator: This tool reviews the live file and overall rule set. It does not currently test whether one specific URL is allowed or blocked for one selected crawler.

Robots.txt Directives You Can Configure

Directive or controlWhat it doesImportant note
User-agentIdentifies the crawler group that should follow the rules beneath it.Use * for all compliant crawlers or select a specific available crawler.
AllowPermits access to a path, often as an exception inside a broader blocked directory.Path matching is case-sensitive.
DisallowAsks the selected crawler not to request a path or directory.Disallow: / applies to the entire site for that group.
SitemapPoints crawlers to an XML sitemap or sitemap index.Use the complete absolute URL, including https:// and the hostname.
Crawl-delayRequests a pause between crawler requests.Googlebot does not support this directive; other crawlers may interpret it differently.
AI-bot controlsCreates crawler-specific rules for AI-bot options available in the interface.Robots.txt relies on voluntary crawler compliance and is not a security barrier.
Directory restrictionsAdds rules for selected folders or URL path patterns.Block only paths you have reviewed; avoid hiding resources needed to render important pages.

Using Wildcards in Robots.txt Paths

Two pattern characters are supported by Google and most major crawlers, and both work in this generator:

PatternMeaningExampleMatches
*Any sequence of charactersDisallow: /*?Any URL containing a query string
$End of the URLDisallow: /*.pdf$Any URL ending in .pdf

Three rules that prevent the most common mistakes:

  • A path with no wildcard is a prefix match. Disallow: /admin blocks /admin/, /admin-tools/, and /administrator/, not just the folder you meant. Add the trailing slash to limit it.
  • When rules conflict, the most specific rule wins for Google, not the first one listed. A longer, more specific Allow overrides a shorter Disallow.
  • Paths are case-sensitive. /Private/ and /private/ are different locations.

Example Robots.txt File

The right file depends on your site structure. This simple example allows general crawling, restricts two directories, and declares a sitemap:

User-agent: *Allow: /Disallow: /admin/Disallow: /private/Sitemap: https://example.com/sitemap.xml

Replace the example paths and domain with real values from your site. Do not publish a template unchanged, and do not list sensitive locations as a substitute for proper access control.

Where to Upload Robots.txt

Publish the file in the top-level directory of the host it controls. For example, rules for https://example.com/ must be available at:

https://example.com/robots.txt

A file at https://example.com/folder/robots.txt does not control the full domain. Rules also apply only to the matching host, protocol, and port. If you use separate subdomains, each one may need its own robots.txt file.

Need a sitemap URL to add to the file? Create one with the XML Sitemap Generator and paste the published sitemap location into the Sitemap field.

Robots.txt Controls Crawling, Not Guaranteed Indexing

Crawling and indexing are different stages. A Disallow rule can stop a compliant crawler from fetching a page, but the URL may still appear in search results when it is discovered through links or other sources. Robots.txt therefore should not be used as the only method for removing a page from search results.

To prevent indexing while allowing a search engine to access the page, use a supported noindex meta tag or X-Robots-Tag. To keep private content inaccessible to crawlers and users, use authentication, permissions, or another server-side control. Do not block a noindex page in robots.txt, because the crawler must fetch the page to see the noindex instruction.

Accuracy rule: A successful validator result means the file appears accessible and the detected rules can be analyzed. It does not guarantee that every allowed URL will be crawled, indexed, ranked, or shown in search results.

Using Robots.txt for AI Bot Crawling Controls

AI companies use named crawlers to gather training data and to fetch pages when answering user questions. Blocking them in robots.txt is currently the standard way to signal that you do not want your content used. This generator includes controls for the ten AI crawlers that matter most in 2026:

CrawlerOperated byWhat it does
GPTBotOpenAICollects data used to train OpenAI models
ChatGPT-UserOpenAIFetches pages on demand when a user asks ChatGPT about them
OAI-SearchBotOpenAIIndexes pages for OpenAI's search features
ClaudeBotAnthropicCollects data used to train Claude
Google-ExtendedGoogleControls use of your content for Gemini and Google's AI training, separate from Googlebot
Applebot-ExtendedAppleControls use of your content for Apple's AI training
CCBotCommon CrawlBuilds the open web archive many AI datasets are derived from
PerplexityBotPerplexityIndexes pages for Perplexity's answer engine
BytespiderByteDanceCollects data for ByteDance and TikTok systems
AmazonbotAmazonCollects data for Amazon services and Alexa

Two things worth understanding before you block everything. 

Training crawlers and answer-engine crawlers do different jobs. Blocking GPTBot stops your content being used to train OpenAI's models. Blocking ChatGPT-User or OAI-SearchBot also stops your pages being retrieved and cited when someone asks a question, which removes you from AI answers entirely. If your goal is visibility in AI search, blocking every AI bot works against you.

Google-Extended is separate from Googlebot. Disallowing Google-Extended does not affect your normal Google Search rankings or crawling. It only controls whether your content feeds Gemini and Google's AI training, so you can block one without the other.

The limitation that applies to all of them: robots.txt is a voluntary protocol. Well-behaved crawlers respect it, but a non-compliant scraper can ignore it entirely. Use server rules, firewall rules, CDN bot management, or authentication when you need enforceable protection rather than a stated preference.

Why Use the Online Tool Pot Generator and Validator?

  • Generate and analyze robots.txt from one tool instead of switching between separate pages.
  • Configure standard crawl rules, sitemap locations, directory restrictions, Crawl-delay, and AI-bot controls.
  • Review live sitemap detection, broad crawler-block status, and crawlability at a glance.
  • Inspect the fetched robots.txt and sitemap information in dedicated result tabs.
  • Copy content directly or download a .txt file for server upload.
  • Add as many rules and restricted directories as your site needs. There is no rule limit.

Related SEO tools: generate an XML sitemap, create structured data, or build a clean SEO slug with other Online Tool Pot utilities.

Common Mistakes to Avoid

❌ Blocking the whole site by accident

✓ Solution:

  • A User-agent group followed by Disallow: / requests a complete crawl block for that group.

❌ Using robots.txt to remove indexed pages

✓ Solution:

  •  Blocked URLs can still appear without a snippet. Use noindex or access control when removal is the goal.

❌ Blocking CSS or JavaScript needed for rendering

✓ Solution:

  • Search engines may struggle to understand pages when essential resources cannot be fetched.

❌ Publishing the file in the wrong folder

✓ Solution:

  • Only the root robots.txt location is recognized for the applicable host.

❌ Entering relative sitemap paths

✓ Solution:

  •  The Sitemap directive should contain a complete absolute URL.

❌ Assuming Crawl-delay works for Google

✓ Solution:

  • Googlebot does not support Crawl-delay, although other crawlers may.

❌ Treating robots.txt as security

✓ Solution:

  • The file is public, and not every bot obeys it. Protect sensitive content at the server or application level.

❌ Ignoring capitalization

✓ Solution:

  • URL paths are case-sensitive. /Private/ and /private/ can match different locations.

Frequently Asked Questions

An online robots txt generator creates a formatted robots.txt file from the crawler rules you select. With this tool, you can configure User-agent, Allow, Disallow, Sitemap, Crawl-delay, AI-bot controls, and directory restrictions, then copy or download the result.

The validator fetches a live robots.txt file and reports sitemap detection, broad administrative blocking, overall crawlability, the retrieved file content, sitemap information, and issues or recommendations detected by the tool.

Not reliably. A blocked URL can still appear in search results if a search engine discovers it elsewhere. Use a supported noindex directive when you want to prevent indexing, and keep the page crawlable so the crawler can read that directive.

Upload it as robots.txt in the root directory of the exact host it controls. For https://example.com/, the file should be available at https://example.com/robots.txt.

No. Googlebot does not support the Crawl-delay directive. Other crawlers may support it and may interpret its value differently.

You can create crawler-specific instructions using the AI-bot controls available in the generator. However, robots.txt depends on voluntary compliance and cannot enforce access against bots that ignore it.

Including one or more absolute Sitemap URLs can help supported crawlers discover your sitemap location. The sitemap line is not tied to one User-agent group.

Yes. You can copy the generated code or download a .txt file and upload it to your server.

No. The current validator analyzes the live robots.txt file and provides a site-level summary. It does not currently run a one-URL, one-user-agent allow-or-block test.

No. You can add as many rules and restricted directories as you need. Keep the file as simple as your site allows, and note that Google only processes the first 500KiB of a robots.txt file, so very large files may be ignored.

Add a rule for GPTBot to stop OpenAI using your content for model training, and rules for ChatGPT-User and OAI-SearchBot to stop your pages being fetched and cited in ChatGPT answers.

No. Google-Extended controls whether your content is used for Gemini and Google's AI training.

Generate or Validate Your Robots.txt File Now

Create crawl rules, add your sitemap, configure search and AI-bot access, and export a server-ready file. Or enter a live URL to review the current robots.txt file, crawlability summary, sitemap status, and recommendations.

Built and maintained by the OnlineToolPot engineering team. Crawler coverage and AI-bot controls were last verified in August 2026.  

More Like This

Calendar Share Link

Calendar Share Link

Create calendar links anyone can add in one click. Works with Google, Outlook, Yahoo, and iCal, with QR codes and recurring event support.

SEO/MarketingBuilt-in
Campaign URL Builder

Campaign URL Builder

Free UTM link builder for GA4. Tag email, social, and paid campaigns with source, medium, campaign, term and content parameters — built one link at a time, with nothing saved.

SEO/MarketingBuilt-in
Email Validator

Email Validator

Free email validator with syntax, MX record, disposable-domain, catch-all and role-based address checks, plus batch validation for whole lists.

SEO/MarketingBuilt-in
Featured alternatives

Five related tools picked to keep users moving.

Calendar Share Link

Top

Campaign URL Builder

Top

Email Validator

Top

Invoice Generator

Top

Advance Loan Calculator

Categories

Dev
AI
Design
Converter/Calculator
SEO/Marketing
Security
Productivity
Text/Writing
Online Tool PotOnline Tool Pot logoOnline Tool Pot dark logo
Login
Online Tool PotToolpot Logo
Cookie Policy|Privacy Policy|Terms of Service|Blogs
© 2026 Online Tool Pot. All rights reserved.