AI Crawler Checker
Enter your site and instantly see whether AI crawlers — GPTBot, ClaudeBot, PerplexityBot, Google-Extended and more — are allowed or blocked by your robots.txt, and whether an llms.txt exists.
Fetches robots.txt and llms.txt through a public CORS proxy. If the fetch fails (some sites block proxies), paste your robots.txt below and hit Check again.
Results
| AI Crawler | Owned by | Status | Rule |
|---|
Which AI crawlers should I allow?
It depends on what you want. Blocking all AI crawlers keeps your content out of AI training runs, but it also means AI assistants and answer engines may never cite or recommend your site. Blocking GPTBot stops OpenAI's training crawler but also affects how ChatGPT browsing answers your site. Google-Extended controls Gemini/AI Overviews usage separately from normal Google Search.
A common strategy in 2026: allow the AI search crawlers (PerplexityBot, OAI-SearchBot, ClaudeBot) so you get cited, and decide individually on training crawlers like GPTBot and CCBot.
How does robots.txt decide who's allowed?
Crawlers follow the most specific group of rules that matches their user-agent name; if no group names them, the * group applies. Within a group, the longest matching Disallow/Allow path wins. This checker implements exactly that logic — the same rules your robots.txt already uses for Googlebot.
Why does llms.txt matter?
An llms.txt file at your site root gives AI models a curated map of your best pages in plain markdown. Robots.txt controls permission; llms.txt provides understanding. Generate yours in seconds with our llms.txt generator, then manage the crawler rules themselves with the robots.txt generator.