Resources
Guides, tutorials, and reference articles about robots.txt and search engine crawling.
Guides
Categories
Knowledge Base
Crawl Budget Explained: Why It Matters for Large Sites
What crawl budget is, how Google allocates crawling resources, why it matters for large sites, and how to optimize your robots.txt and site architecture to make the most of it.
Read moreDo You Need a robots.txt File?
Does your website need a robots.txt file? When it's essential, when it's optional, and what happens if you don't have one.
Read moreHow Googlebot Works: The Complete Guide
How Googlebot discovers, crawls, and renders pages. The different Googlebot user agents, crawl frequency, JavaScript rendering, robots.txt interaction, and how to verify real Googlebot.
Read moreHow to Test If Googlebot Can Access Your Pages
How to test whether Googlebot can access, crawl, and render your pages. Covers Google Search Console tools, robots.txt testing, fetch and render, and common access issues.
Read moreHow Search Engines Actually Crawl Your Site
A practical explanation of how search engine crawlers discover, fetch, render, and index web pages. Covers Googlebot, Bingbot, crawl scheduling, rendering, and what affects crawl behavior.
Read moreHow to Add Your Sitemap to robots.txt
Add your sitemap URL to robots.txt so search engines can find it. Syntax, multiple sitemaps, and common mistakes to avoid.
Read moreHow to Block AI Crawlers with robots.txt
Block AI training crawlers like GPTBot, ClaudeBot, and PerplexityBot using robots.txt. Complete list of AI user agents and copy-paste rules.
Read moreHow to Check Any Website's robots.txt File
How to find and check any website's robots.txt file. View directives, verify accessibility, and check for common configuration mistakes.
Read moreHow to Create a robots.txt File
Step-by-step guide to creating a robots.txt file. Learn the syntax, write directives for search engines, and deploy your file correctly.
Read moreHow to Edit robots.txt in Shopify
How to customize your Shopify store's robots.txt using the robots.txt.liquid template. Default rules, common customizations, and gotchas.
Read moreHow to Edit robots.txt in WordPress
Edit your WordPress robots.txt file using Yoast SEO, Rank Math, or direct file editing. Step-by-step instructions for each method.
Read moreHow to Fix "Blocked by robots.txt" Errors
Fix 'blocked by robots.txt' errors in Google Search Console. Diagnose Disallow rules, update your robots.txt, and get your pages indexed.
Read moreHow to Read and Understand robots.txt
Learn to read any robots.txt file. Understand User-agent, Disallow, Allow, Sitemap, and Crawl-delay directives with real-world examples.
Read moreHow to Test and Validate Your robots.txt
Four ways to test your robots.txt file: online validators, Google Search Console, command-line tools, and dedicated testing services.
Read morellms.txt vs robots.txt: How They Work Together
How llms.txt and robots.txt complement each other for AI content access. What each file does, comparison table, and how to use both together.
Read moreWhat Is the Robots Exclusion Protocol?
The history and mechanics of the Robots Exclusion Protocol, from its 1994 origins to RFC 9309. How crawlers implement it, why it is advisory, and what it means for the modern web.
Read morerobots.txt Allow Directive: How to Override Disallow Rules
How the Allow directive works in robots.txt, how it interacts with Disallow, specificity rules, practical examples, and how to test Allow rules.
Read moreHow robots.txt Affects Your SEO
How robots.txt impacts search engine optimization. Crawl budget, indexing, and the SEO mistakes that robots.txt can cause.
Read morerobots.txt Best Practices
Robots.txt best practices for SEO and crawl management. What to block, what to allow, and the mistakes that hurt your site.
Read moreCrawl-Delay in robots.txt Explained
What the Crawl-delay directive does, which search engines support it, and when you should (and shouldn't) use it.
Read morerobots.txt Glossary: Every Directive and Term Explained
Definitions of every robots.txt directive and related term. User-agent, Disallow, Allow, Sitemap, Crawl-delay, and more.
Read morerobots.txt Disallow Directive Explained
How the Disallow directive works in robots.txt. Syntax, examples, path matching, and common mistakes that accidentally block your entire site.
Read morerobots.txt for Drupal Sites
How to configure robots.txt for Drupal sites. Covers the default file, customization methods, Drupal-specific paths to block, and common issues with modules and updates.
Read morerobots.txt Examples for Every Situation
Copy-paste robots.txt examples: block all bots, allow all bots, block specific directories, WordPress, Shopify, e-commerce, and more.
Read moreWhy You Should Monitor Your robots.txt for Changes
Accidental robots.txt changes can deindex your entire site. Why monitoring your robots.txt matters and how to set it up.
Read moreNoindex in robots.txt: Why It Doesn't Work
Google no longer supports the noindex directive in robots.txt. What happened, what to use instead, and how to properly deindex pages.
Read moreDoes robots.txt Prevent Indexing? (No, and Here's Why)
Why blocking a URL in robots.txt does not prevent it from appearing in Google search results. The difference between crawling and indexing, and what to use instead.
Read morerobots.txt SEO Audit Checklist
A checklist for auditing your robots.txt file. Check for SEO issues, crawl problems, and misconfigurations in under 10 minutes.
Read morerobots.txt for Squarespace Sites
How robots.txt works on Squarespace, what the default file contains, what you can and cannot customize, and how to handle common Squarespace crawling issues.
Read morerobots.txt for Staging and Development Sites
How to prevent staging and development sites from getting indexed by Google. Covers robots.txt, noindex headers, password protection, and environment-specific configuration.
Read morerobots.txt Syntax Reference
Complete robots.txt syntax reference. Every directive, pattern, and rule explained with examples. Bookmark this page.
Read morerobots.txt User-Agent: How to Target Specific Crawlers
How to use the User-agent directive in robots.txt to create rules for specific search engines, bots, and crawlers.
Read morerobots.txt vs Meta Robots Tags: Which to Use
The difference between robots.txt and meta robots tags (noindex, nofollow). When to use each, and why using the wrong one can hurt your SEO.
Read morerobots.txt for Webflow Sites
How to configure robots.txt on Webflow. Covers the auto-generated file, custom editing, common configurations, and Webflow-specific crawling considerations.
Read moreWildcards in robots.txt: Using * and $ Patterns
How to use wildcard patterns in robots.txt. The * and $ characters, path matching, and practical examples for complex blocking rules.
Read morerobots.txt for Wix, Squarespace, and Webflow
How Wix, Squarespace, and Webflow handle robots.txt. Default rules, customization options, workarounds, and platform-specific gotchas for each website builder.
Read moreComplete List of Search Engine and AI Bot User-Agents
A comprehensive reference of search engine crawler and AI bot user-agent strings. Covers Google, Bing, Yandex, Baidu, DuckDuckGo, and major AI crawlers with their robots.txt identifiers.
Read moreWhat Does robots.txt Actually Do?
What robots.txt does and doesn't do. How crawlers use it, why it's advisory not enforceable, and the limits of robots.txt.
Read moreWhat Is llms.txt? The New Standard for AI Content Access
What llms.txt is, how it works, how it differs from robots.txt, the file format, current adoption, and how to create one for your site.
Read moreWhat Is robots.txt?
What robots.txt is, how it works, and why every website should have one. The complete introduction to the robots exclusion protocol.
Read moreWhat Is Web Crawling? How Crawlers Work
A plain-English explanation of web crawling: what it is, how web crawlers work, how they discover and process pages, and how robots.txt controls their behavior.
Read morerobots.txt vs X-Robots-Tag: The HTTP Header Approach
What the X-Robots-Tag HTTP header is, how it differs from robots.txt and meta robots, when to use each, and how to implement X-Robots-Tag on Nginx, Apache, and CDNs.
Read more