robots.txt
robots.txt is a text file at the root of a site that tells crawlers which sections they may or may not crawl.
Check whether a URL is allowed or blocked for a given crawler (RFC 9309).
robots.txt is a text file at the root of a site that tells crawlers which sections they may or may not crawl.
Crawl budget is the number of URLs a crawler can and wants to crawl on a site in a given period.
Tell me about your project
Leave your site’s address and goal — within one business day you’ll get first observations and next steps.
Or write directly
Paste a site or page URL — we load its robots.txt and check the page right away.
Allowed
No rule matches
Matching follows RFC 9309: longest rule wins, Allow wins ties, * and $ supported.
Declared sitemaps