AI Crawlers & Bots Directory
Twelve crawlers decide how AI systems see your website — from OpenAI's GPTBot and Anthropic's ClaudeBot to PerplexityBot and Bingbot. Each page explains what the bot does, what happens if you block it, and the exact robots.txt rule that controls it.
OpenAI's crawler that collects publicly available web pages to train its generative AI models. It is separate from ChatGPT-User (real-time fetches) and OAI-SearchBot (search indexing).
Read the guide →OpenAI's search crawler. It indexes the pages that power ChatGPT's search feature and the source links ChatGPT shows next to answers.
Read the guide →Fetches pages on behalf of an individual ChatGPT user at their request — for example when ChatGPT opens a link you pasted into the conversation. It is not a bulk crawler.
Read the guide →Anthropic's crawler that gathers public web data for training and evaluating its Claude models. Anthropic identifies it through its user-agent token and controls it via robots.txt.
Read the guide →Anthropic's crawler for improving and grounding the sources Claude surfaces when it answers questions with web results.
Read the guide →Fetches pages in real time when an individual Claude user asks it to open or read a specific URL during a conversation.
Read the guide →Perplexity's crawler that maintains the search index behind its AI answer engine. Perplexity documents it as controllable through robots.txt.
Read the guide →Not a standalone search crawler — it is a controls token you add to robots.txt to manage whether Google may use your content for Gemini apps and Vertex AI grounding, independently of Google Search. Appearing in Google Search (and AI Overviews) follows Googlebot rules, not Google-Extended.
Read the guide →Google Cloud's crawler that fetches content for Vertex AI Search services and customer-built search deployments on Google Cloud.
Read the guide →Microsoft's search crawler. Bing's index built by Bingbot also feeds Bing chat and Copilot answers, so it is the Bing equivalent of Googlebot.
Read the guide →Microsoft's rendering crawler that takes snapshots of pages so Bing can show link previews and rich snippets that match the rendered page.
Read the guide →A controls token for Apple's robots.txt: it governs whether your content may be used to train Apple's foundation models behind Apple Intelligence and Siri. General Applebot (Apple's search crawler) is a separate user-agent.
Read the guide →Check your own robots.txt
Not sure whether your site allows these crawlers? Run your robots.txt through the checker — it flags missing or conflicting AI bot rules.
Open the Robots.txt Checker