Tools · SEO Tools
AI Crawler Checker
See how a robots.txt file treats known AI crawlers. robots.txt is a preference — not a hard guarantee.
Registry
- GPTBot (OpenAI) — May crawl content that could be used in training generative AI foundation models. Source
- OAI-SearchBot (OpenAI) — Surfaces websites in ChatGPT search features. Source
- ChatGPT-User (OpenAI) — User-initiated fetches in ChatGPT; robots.txt rules may not apply. Source
- ClaudeBot (Anthropic) — Collects web content that could contribute to model training. Source
- Claude-SearchBot (Anthropic) — Indexes content to improve Claude search result quality. Source
- Claude-User (Anthropic) — User-directed retrieval when people ask Claude questions. Source
- Google-Extended (Google) — robots.txt control token for Gemini training/grounding use of crawled content; not a separate HTTP crawler UA. Source
- Applebot-Extended (Apple) — robots.txt control for Apple foundation-model training use of Applebot-crawled content; does not crawl separately. Source
- PerplexityBot (Perplexity) — Surfaces and links websites in Perplexity search results; documented as not for foundation-model training crawl. Source
- Perplexity-User (Perplexity) — User-initiated fetches; Perplexity documents that this fetcher generally ignores robots.txt. Source
Compliant crawlers may honor robots.txt; others may not. This is not access control. API: None.
Read the original guide: Compare robots.txt to a verified AI crawler registry.