GuideAI crawlers

03User-agent reference

AI crawler user-agent reference

14 documented AI crawler user-agents: operator, answer engine, and exactly what allowing or blocking each one does to your AI visibility. Use this when auditing your robots.txt.

Last updated

01High-impact crawlers

Block these and your AI visibility drops significantly.

These crawlers feed the main AI answer engines that most buyers use today: ChatGPT, Google AI, Perplexity, and Copilot. Blocking any of them in robots.txt removes you from the corresponding engine.

OAI-SearchBot

OpenAI

High impact
ChatGPT (web search)
Powers ChatGPT's real-time web search feature. When a user asks ChatGPT a question with web search enabled, OAI-SearchBot may fetch pages to support the answer.
If allowed: ChatGPT can fetch and cite your pages in real-time web search responses. Affects ChatGPT's ability to surface your content in search-grounded answers.
If blocked: ChatGPT cannot directly fetch your pages for search-grounded answers. However, ChatGPT also retrieves via Bing — pages indexed in Bing may still appear via Bing-backed retrieval.
ClaudeBot

Anthropic

High impact
Claude AI
Primary crawler for Anthropic's Claude AI. Crawls pages to build the knowledge base and context Claude uses when grounding answers about websites and brands.
If allowed: Claude AI can access and reference your site's content. Improves Claude's ability to describe, summarise, and cite your brand accurately.
If blocked: Claude cannot crawl your site. Claude's knowledge of your brand and content will be limited to what was in training data or what users paste in — no real-time grounding from your site.
PerplexityBot

Perplexity

High impact
Perplexity AI
Perplexity's crawler. Perplexity runs live search at query time (Google + Bing), but also uses its own crawl index for certain content. PerplexityBot contributes to that index.
If allowed: Your content can be directly indexed and cited by Perplexity AI answers. Combined with Bing and Google indexing, this maximises Perplexity visibility.
If blocked: Perplexity cannot directly crawl your pages. Perplexity may still surface your content via Bing or Google results it queries at runtime, but direct crawl-based indexing is prevented.
Googlebot

Google

High impact
Google AI OverviewsGoogle AI ModeGeminiGoogle Search
Google's primary web crawler. Powers all Google search and all Google AI surfaces. Blocking Googlebot removes your content from Google Search, Google AI Overviews, Google AI Mode, and Gemini grounding.
If allowed: Your pages are eligible for Google Search indexing and all Google AI surfaces: AI Overviews, AI Mode, and Gemini grounding. Essential for any Google AI visibility.
If blocked: CRITICAL: your pages are removed from Google Search and all Google AI surfaces. This is the highest-impact block possible. Only block Googlebot for content you explicitly want excluded from all Google surfaces.
Bingbot

Microsoft

High impact
ChatGPT (via Bing)Microsoft CopilotBing Search
Microsoft's primary web crawler. Powers Bing Search and is the retrieval backbone for both ChatGPT web search (87% of ChatGPT citations match Bing top-10) and Microsoft Copilot.
If allowed: Your pages are eligible for Bing Search indexing, which directly feeds ChatGPT web browsing and Copilot grounding. Blocking Bingbot is equivalent to blocking ChatGPT and Copilot visibility.
If blocked: CRITICAL: your pages are removed from Bing Search and from ChatGPT/Copilot AI grounding. Since ~87% of ChatGPT citations match Bing top-10, a Bingbot block is highly impactful.

02Medium-impact crawlers

Allow these to broaden AI reach — blocking them has measurable effect.

These crawlers contribute to real-time or supplementary AI retrieval. Blocking them reduces AI access in specific contexts but does not remove you from the primary index.

GPTBot

OpenAI

Medium impact
ChatGPT (training data)
Crawls the open web for OpenAI model training data. Blocking it opts your content out of future OpenAI training runs.
If allowed: Your content may be included in future OpenAI training datasets, potentially improving how well ChatGPT understands and describes your brand.
If blocked: Your content is excluded from OpenAI training data. Does NOT affect ChatGPT's real-time web browsing (that uses OAI-SearchBot) or search results (Bing-powered).
ChatGPT-User

OpenAI

Medium impact
ChatGPT (browsing)
Used by ChatGPT when it browses the web during a user conversation — the "browse" or "search" capability in the assistant. Fetches specific URLs in real time.
If allowed: ChatGPT can visit and read specific URLs during conversations where the user or ChatGPT requests a page.
If blocked: ChatGPT cannot browse to your pages directly during conversations. The user can still paste content manually; Bing-indexed results may still be retrieved.
Claude-Web

Anthropic

Medium impact
Claude AI (web access)
Used by Claude when performing real-time web browsing within conversations. Allows Claude to fetch specific URLs a user references or that Claude decides to look up.
If allowed: Claude can browse to your pages in real time during conversations, including fetching specific URLs for summarisation or citation.
If blocked: Claude cannot browse to your pages during conversations. Real-time browsing to your site is blocked.
Amazonbot

Amazon

Medium impact
AlexaAmazon Rufus (Amazon AI shopping)
Amazon's crawler for Alexa and Amazon AI features including Rufus, Amazon's AI shopping assistant. Rufus uses web-crawled product and brand content to answer shopping queries.
If allowed: Amazon AI features (Rufus, Alexa) can index and reference your brand and product content for shopping-related queries.
If blocked: Amazon AI cannot directly crawl your pages. Brand and product discoverability in Amazon AI features is reduced.

03Low-impact crawlers for current AI answer engines

Training data and emerging surfaces — allow unless you have a specific reason not to.

These crawlers primarily feed AI training datasets or emerging AI surfaces not yet measured by Lokrix AI. Blocking them has minimal near-term effect on your ChatGPT / Google / Perplexity / Gemini scores, but may affect future model generations.

Google-Extended

Google

Low impact
Google AI (training data only)
An opt-out user-agent for Google AI training and Bard/Gemini improvement data. IMPORTANT: blocking Google-Extended does NOT block Googlebot, AI Overviews, or Gemini grounding — it only opts out of AI training datasets.
Allow: Your content may be used in Google AI model training and improvement (Bard, Gemini, etc.).
Applebot-Extended

Apple

Low impact
Apple IntelligenceSiri
Apple's extended crawler for Apple Intelligence and Siri. Similar to Google-Extended in structure — an opt-out signal for Apple AI training data, separate from the main Applebot (which powers Spotlight and Siri web results).
Allow: Your content may be used for Apple Intelligence and Siri AI features, including potential grounding for Siri web-based answers.
CCBot

Common Crawl

Low impact
Open-source LLMs (via training data)Various AI models
Common Crawl's crawler. Common Crawl is a nonprofit that provides open web crawl data widely used in AI model training (GPT-3/4, Llama, Mistral, and many others were trained on CC datasets).
Allow: Your content enters the Common Crawl corpus, which is used as training data by many open-source and commercial AI models. Indirect but broad AI model awareness of your brand.
Bytespider

ByteDance (TikTok)

Low impact
TikTok AI featuresByteDance AI products
ByteDance's crawler, used for TikTok AI features and ByteDance AI products. Primarily relevant for brands targeting TikTok-native AI search and recommendation features.
Allow: Your content can be indexed for TikTok AI features and ByteDance AI recommendations.
Meta-ExternalAgent

Meta

Low impact
Meta AI (Facebook, Instagram, WhatsApp)Llama (training data)
Meta's crawler for Meta AI features across Facebook, Instagram, and WhatsApp, and for Llama model training data collection.
Allow: Your content can be indexed for Meta AI features (Meta AI assistant in Facebook/Instagram/WhatsApp) and may contribute to Llama training datasets.

04Quick reference

All 14 crawlers at a glance.

GPTBot
OpenAI
ChatGPT (training data)
Medium impact
OAI-SearchBot
OpenAI
ChatGPT (web search)
High impact
ChatGPT-User
OpenAI
ChatGPT (browsing)
Medium impact
ClaudeBot
Anthropic
Claude AI
High impact
Claude-Web
Anthropic
Claude AI (web access)
Medium impact
PerplexityBot
Perplexity
Perplexity AI
High impact
Google-Extended
Google
Google AI (training data only)
Low impact
Googlebot
Google
Google AI OverviewsGoogle AI ModeGeminiGoogle Search
High impact
Bingbot
Microsoft
ChatGPT (via Bing)Microsoft CopilotBing Search
High impact
Amazonbot
Amazon
AlexaAmazon Rufus (Amazon AI shopping)
Medium impact
Applebot-Extended
Apple
Apple IntelligenceSiri
Low impact
CCBot
Common Crawl
Open-source LLMs (via training data)Various AI models
Low impact
Bytespider
ByteDance (TikTok)
TikTok AI featuresByteDance AI products
Low impact
Meta-ExternalAgent
Meta
Meta AI (Facebook, Instagram, WhatsApp)Llama (training data)
Low impact

05Engine grounding cross-reference

How crawlers map to retrieval indexes and answer engines.

The same content can be indexed by multiple crawlers — but only the ones matching an engine's retrieval index affect that engine's answers. Source: Answer-engine citation-source analyses (2024–2026), summarised in CLAUDE.md §4.

ChatGPT
Bing
Bing top-10 rank + freshness(87% citation index match)
Google AI Overviews
Google
Google top-20; separate from AI Mode(54% citation index match)
Google AI Mode
Google
Shares a cited URL with AIO only ~14% of the time
Perplexity
Google + Bing (live)
Over-weights Reddit (~47% of top citations) + freshness(47% citation index match)
Gemini
Google
Google Search grounding; schema-sensitive
Grok (xAI)
X + live web
Grounds on X posts + live web; freshness-weighted (beta — live via a web-search fallback; draws not yet counted in headline scores)
Claude
Parametric + ClaudeBot
Factual density and entity clarity are the primary levers; parametric knowledge with shallow ClaudeBot live search

Source: Answer-engine citation-source analyses (2024–2026), summarised in CLAUDE.md §4.

06robots.txt template

Allow all high-impact AI crawlers with one block.

The minimal robots.txt configuration to ensure all high-impact AI crawlers can access your site. Adjust the Sitemap URL to match your domain.

robots.txt — allow all AI crawlers
# Allow all crawlers (default open)
User-agent: *
Allow: /

# Explicitly allow high-impact AI crawlers
# (redundant if User-agent: * is already open, but explicit for clarity)
User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Claude-Web
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Googlebot
Allow: /

# Google-Extended: controls AI training data, NOT AI Overviews/Gemini
# Allow to permit Google AI training; Disallow to opt out
User-agent: Google-Extended
Allow: /

User-agent: Bingbot
Allow: /

# Sitemap — helps all crawlers discover your pages
Sitemap: https://yourdomain.com/sitemap.xml

Adjust Sitemap URL. Add specific Disallow rules only for pages you explicitly want excluded (login pages, admin areas, etc.).

Crawlers allowed. Now measure your AI visibility.

Lokrix AI measures your presence probability across ChatGPT, Google AI Overviews, AI Mode, Perplexity, Gemini, Grok (beta), and Claude — with a 95% CI on every score.

Start free · no credit card