The crawler registry

AI crawler directory

These records drive the policy matrix. Robots tokens are central; full User-Agent browser versions may change.

ChatGPT-User

OpenAI · assistant

User-initiated retrieval; robots.txt may not apply. This is not the ChatGPT search-index crawler.

Robots token: ChatGPT-User

Source reviewed

OpenAI crawler documentation ↗

ClaudeBot

Anthropic · training

Collects material that may contribute to model training; an independent policy choice.

Robots token: ClaudeBot

Source reviewed

Anthropic crawler documentation ↗

Claude-User

Anthropic · assistant

Retrieves pages in response to user requests. Anthropic documents robots.txt compliance.

Robots token: Claude-User

Source reviewed

Anthropic crawler documentation ↗

Perplexity-User

Perplexity · assistant

User-requested retrieval generally ignores robots.txt. Policy here is not proof of enforcement.

Robots token: Perplexity-User

Source reviewed

Perplexity crawler documentation ↗

Googlebot

Google · search

Google Search crawler. Google-Extended is a distinct use-control token.

Robots token: Googlebot

Source reviewed

Google crawler documentation ↗

bingbot

Microsoft · search

Bing Search crawler. Read robots policy independently from verified crawler identity.

Robots token: bingbot

Source reviewed

Microsoft crawler documentation ↗

Applebot

Apple · mixed

Powers Apple discovery and other documented uses. Applebot-Extended separately controls training use; Googlebot policy is the documented fallback.

Robots token: Applebot

Source reviewed

Apple crawler documentation ↗

CCBot

Common Crawl · mixed

Builds a public web corpus usable for research and other downstream purposes. Blocking it is not a search-readiness failure.

Robots token: CCBot

Source reviewed

Common Crawl crawler documentation ↗

Amazonbot

Amazon · mixed

Collects public content for Amazon services. Consult the operator for current use and control details.

Robots token: Amazonbot

Source reviewed

Amazon crawler documentation ↗