List of All AI Crawlers and Bots
What Are AI Crawlers and Bots (and the Four Types That Matter)?
An AI crawler or bot is an automated program a company runs to fetch web pages, but “AI crawler” covers four different jobs: standing crawlers that gather training data, AI-specific crawlers that build live search or retrieval indexes, on-demand bots that fetch a page because a person asked for it, and traditional search-engine crawlers whose indexes are also used to ground AI-generated answers. Standing bots are crawlers that run independently of a specific user request.
Methodology
This article’s fact base comes primarily from each vendor’s own documentation: OpenAI, Anthropic, Perplexity, Google, Apple, Meta and Mistral publish dedicated bot or robots pages, and Common Crawl and DuckDuckGo do the same for their crawlers.
Training Crawlers: The Bots That Feed Model Development
Training crawlers are the standing crawlers whose stated job is gathering content for model training, not search or one-off fetches: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, MistralAI-Training, Amazonbot and CCBot, each documented by its operator with its own distinct user-agent token.
| Bot | Operator | Category | Purpose | User-Agent |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | ”crawl content that may be used in training our generative AI foundation models” | GPTBot (e.g. GPTBot/1.4) |
| ClaudeBot | Anthropic | Training | Collects web content to “enhance the utility and safety of” Anthropic’s generative AI models | ClaudeBot |
| Google-Extended | Training | Opt-out token for whether Google-crawled content trains future Gemini models and grounds Gemini Apps/Vertex AI; no effect on Search ranking | Google-Extended | |
| Applebot-Extended | Apple | Training | Control token governing whether content Applebot already fetched may train Apple’s generative AI models; does not crawl on its own | Applebot-Extended |
| Meta-ExternalAgent | Meta | Training | Crawls for “training foundation AI models or improving products by indexing content directly” | meta-externalagent/1.1 |
| MistralAI-Training | Mistral | Training | Crawls web content “to help build datasets for training Mistral generative AI models” | MistralAI-Training/1.0 |
| Amazonbot | Amazon | Training | Improves Amazon products including Alexa; collected content may train Amazon’s generative AI models | Amazonbot/0.1 |
| CCBot | Common Crawl | Training | Builds Common Crawl’s open web archive, a dataset widely reused as AI training input across the industry | CCBot/2.0 |
One pattern is worth calling out: most of these operators back the user-agent string with a verifiable IP range, not just a name you could spoof. GPTBot’s IP ranges are published at openai.com/gptbot.json, so a request claiming to be GPTBot can be confirmed as OpenAI’s, not a copycat. Common Crawl does the same, and its own documentation is unusually candid about the risk: it warns that other crawlers spoof the CCBot user-agent string, and recommends reverse-DNS verification against its published IP list rather than trusting the string alone.
Applebot-Extended is the odd row in this table. It isn’t a crawler; it’s a permission flag governing reuse of pages Applebot already fetched. The Amazonbot row is also worth flagging now: I couldn’t independently re-fetch Amazon’s own developer page this research pass, so treat its operator details and user-agent format as lower-confidence than the rest of the table.
Standing Search and Retrieval-Index Crawlers
Standing search and retrieval-index crawlers build the index these companies use to answer live questions and surface citations, explicitly not for training: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer, MistralAI-Index, DuckAssistBot and Apple’s narrow podcast crawler iTMS. Each vendor documents this class as walled off from the training crawlers in the previous section, with its own purpose statement and user-agent token, not a relabeled version of the same bot.
| Bot | Operator | Category | Purpose | User-Agent |
|---|---|---|---|---|
| OAI-SearchBot | OpenAI | Live web search | Powers ChatGPT’s search and citation features; not used for model training | OAI-SearchBot (e.g. OAI-SearchBot/1.4) |
| Claude-SearchBot | Anthropic | Live web search | ”navigates the web to improve search result quality” for Claude | Claude-SearchBot |
| PerplexityBot | Perplexity | Live web search | ”designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models.” | PerplexityBot/1.0 |
| Meta-WebIndexer | Meta | Live web search | Improves Meta AI’s search result quality through content analysis | meta-webindexer/1.1 |
| MistralAI-Index | Mistral | Live web search | Automated indexing for Le Chat search, “not used for generative AI training” | MistralAI-Index/1.0 |
| DuckAssistBot | DuckDuckGo | Live web search | Crawls pages in real time for AI-assisted answers that cite sources | DuckAssistBot/1.2 |
| iTMS | Apple | Live web search | ”only crawls URLs associated with registered content on Apple Podcasts” | iTMS |
Search Engine Crawlers Used by AI Systems
Some AI products do not rely only on dedicated AI crawlers. They also draw from traditional search-engine indexes. These crawlers primarily exist to build general web-search indexes, but those same indexes can also be used to ground AI-generated answers.
| Bot | Operator | Category | Purpose | User-Agent |
|---|---|---|---|---|
| Googlebot | Search index used by AI | Builds Google Search’s web index, which also supports AI search experiences such as AI Overviews and AI Mode. | Googlebot | |
| bingbot | Microsoft | Search index used by AI | Builds Bing’s web index, which is also used for web-grounded Microsoft Copilot experiences. | bingbot |
| Bravebot | Brave | Search index used by AI | Builds Brave Search’s independent web index, which can also be accessed by AI applications through Brave Search.Claude can also use Brave to retrieve web results for grounded answers. | Bravebot (robots.txt token) |
This category is different from AI-specific search crawlers such as OAI-SearchBot or PerplexityBot. Googlebot, bingbot and Brave’s crawler are general-purpose search crawlers first; their relevance to AI comes from the fact that the indexes they build are also used by AI systems.
For Google and Microsoft, that means blocking the underlying search crawler can affect both traditional search visibility and AI experiences that depend on the same index. Brave is slightly different operationally: Brave documents Bravebot as the robots.txt token for controlling its search crawler, while the crawler itself does not necessarily expose a distinct Brave-specific HTTP user-agent on every request.
On-Demand User-Fetch Bots: Only Active When Someone Asks
On-demand user-fetch bots aren’t really crawlers at all. They fire once, on request, when a real person pastes a URL or asks an assistant about a specific page: ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher, MistralAI-User, Google-CloudVertexBot and OAI-AdsBot. That user-triggered pattern matters for robots.txt: because a person, not the bot, decided the page mattered, several of these fetches don’t behave like standing crawlers at all.
| Bot | Operator | Category | Trigger | User-Agent |
|---|---|---|---|---|
| ChatGPT-User | OpenAI | On-demand user fetch | Fires when a person pastes a URL into ChatGPT or triggers a Custom GPT/GPT Action | ChatGPT-User (e.g. ChatGPT-User/1.0) |
| Claude-User | Anthropic | On-demand user fetch | Fires when a Claude user’s question requires retrieving a specific web page | Claude-User |
| Perplexity-User | Perplexity | On-demand user fetch | Fires “when users ask Perplexity a question” and it needs a specific page to answer | Perplexity-User/1.0 |
| Meta-ExternalFetcher | Meta | On-demand user fetch | Fetches specific links at user request, including agentic AI features | meta-externalfetcher/1.1 |
| MistralAI-User | Mistral | On-demand user fetch | Fires only when a Le Chat user’s question needs a specific page | MistralAI-User/1.0 |
| Google-CloudVertexBot | On-demand user fetch | Crawls only when a site owner directly requests it, to build a Vertex AI Agent | Google-CloudVertexBot | |
| OAI-AdsBot | OpenAI | On-demand user fetch | Validates the safety of web pages submitted as ads on ChatGPT | OAI-AdsBot (e.g. OAI-AdsBot/1.0) |
Two rows in this table work a little differently from the rest. OAI-AdsBot isn’t triggered by a chat question at all. It validates the safety of pages submitted as ads on ChatGPT, a narrower and more commercial job than the other six. Google-CloudVertexBot is the opposite kind of exception: it doesn’t even wait for an end user’s question. It fires only when a site owner directly requests it, to build a Vertex AI Agent, an opt-in relationship none of the other on-demand bots have. The rest, ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher and MistralAI-User, share the same trigger: a person types a URL or asks a question, and the assistant fetches exactly that one page, nothing more.
Because the fetch is user-initiated rather than automated crawling, some of these bots don’t behave like standing crawlers at all. Meta’s own documentation says Meta-ExternalFetcher “may bypass robots.txt rules” for these user-requested fetches. Perplexity’s docs make a similar point about Perplexity-User: it generally ignores robots.txt because a person, not the bot, decided the page mattered. That distinction matters directly for robots.txt decisions: a rule that stops a standing crawler doesn’t necessarily stop a user-triggered fetch.
Bots Without Official Documentation: Bytespider and Grok
Bytespider and Grok are two widely-discussed AI bots with no official vendor documentation: ByteDance’s Bytespider and xAI’s Grok crawler, neither with a confirmed purpose statement or verified user-agent format. That absence has a practical cost: without a published IP list or purpose statement, there’s no way to verify that a request claiming one of these names is actually the vendor, the way GPTBot’s or CCBot’s published IP ranges allow elsewhere in this article.
Which AI Bots Should You Allow or Block in robots.txt?
If your goal is maximum visibility in AI search, allow all relevant AI crawlers. Rather than blocking an entire crawler, block only specific pages or directories you don’t want accessed. You may also choose to limit crawlers if extensive crawling creates meaningful server costs.
What you should allow depends on your goal:
- Live AI search visibility: Make sure search and retrieval crawlers can access your content.
- Future visibility in trained models: Make sure training crawlers are allowed.
- Maximum AI visibility: Allow both.
For on-demand user-fetch bots, robots.txt controls may work differently because the fetch is triggered by a person rather than a standing crawler. When blocking a crawler, match the stable bot name rather than a versioned user-agent string.
User-agent: ClaudeBot
Disallow: /private/
Cloudflare’s Content Signals Policy adds more granular preferences such as search, ai-input, and ai-train, allowing site owners to distinguish traditional search, real-time AI use, and model training.
Why AI crawlers and bots matter
Crawl access is becoming a negotiated, policy-governed layer of the web. On July 1, 2025, Cloudflare said it was “changing the default to block AI crawlers unless they pay creators for their content” for newly onboarding domains. Cloudflare protects roughly a fifth of all web traffic overall, and that new pay-for-crawl default now covers a large share of the sites newly onboarding to it, a narrower claim than the network-wide figure.
Whether these bots visit your pages at all determines whether you can appear in an AI answer in the first place. That’s the practical stakes behind every AI crawler and bot covered in this article: no crawl, no citation, no matter how good the page is. Ranking well in Google doesn’t guarantee AI crawlers show up on the same schedule, since each one runs its own crawl queue.
Here’s the honesty point most articles skip. Crawling is passive, and its timing isn’t controlled by the site owner. Most site owners only learn whether GPTBot, ClaudeBot, PerplexityBot or any of the others actually visited by grepping server logs, after the fact. There’s no dashboard notification, no “your page was just indexed” email. You find out you were ignored the same way you find out you were crawled: by checking.
How Do You Get AI Bots to Crawl Your Content Faster?
ALLMO’s Warm-Up answers this by actively triggering AI crawler visits from ChatGPT and Perplexity instead of waiting for passive discovery, and it brings crawlers back over several days rather than a single pass. Every AI crawler and bot covered in this article is otherwise either on a fixed schedule or waiting for a user to ask about your specific page, which means a new or updated page can sit unvisited for days.
Warm-Up doesn’t guarantee citations. It accelerates discovery and processing, closing the gap between publishing a page and a bot actually fetching it. Those are different problems, and it’s worth being precise about which one you’re solving.
The case study is the clearest evidence I have for the gap itself. When aicitationsgenerator.com launched, all 6 submitted pages were crawled by ChatGPT and Perplexity within 10 minutes of triggering Warm-Up, with the first crawl landing at 3 minutes, a 100% success rate. Without Warm-Up, 77% of pages went uncrawled.
That 77% is the number that matters most, more than knowing any single crawler’s name or user-agent string. These bots rarely refuse to crawl a page outright. Most pages are simply never on their radar in time to matter. If you want to see what triggering that visit looks like on your own site, ALLMO’s Warm-Up is the feature built for exactly this gap.
Frequently Asked Questions
GPTBot vs. PerplexityBot vs. ClaudeBot: what's the difference?
GPTBot and ClaudeBot are standing training crawlers: OpenAI and Anthropic each document them as collecting web content specifically to train their generative AI models. PerplexityBot is different, Perplexity states it only surfaces and links websites in Perplexity's search results and is not used to crawl content for AI foundation models. Block GPTBot and ClaudeBot to opt out of training; PerplexityBot doesn't feed training at all.
How do I know if AI crawlers are actually indexing my site?
Crawling and indexing are different events, and most AI labs don't send a notification for either. The only reliable check is grepping your server logs for user-agent tokens like GPTBot, ClaudeBot or PerplexityBot to confirm a visit happened. A log entry proves the bot fetched the page, not that the page was used in an AI answer.
Does blocking a bot in robots.txt actually stop it from crawling?
For standing crawlers that honor it, yes, robots.txt blocking works as intended, which covers most major labs in this article. But the rule trusts the user-agent string, not identity: anyone can send a 'GPTBot' or 'CCBot' header without being OpenAI or Common Crawl. That's why OpenAI publishes GPTBot's IP ranges (openai.com/gptbot.json) and Common Crawl recommends reverse-DNS checks, since its own docs warn other crawlers spoof the CCBot string. Robots.txt alone can't catch that.
Track your AI search visibility
See where your brand appears in ChatGPT, Perplexity, and other AI search engines.