All articles
AI Platforms

AI Crawler Robots.txt Guide: Every Bot You Need to Allow (2026 Update)

13 AI crawlers now scan the web — OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Meta-ExternalAgent, MistralAI-User, and more. Rebuilt August 2026 on SEOmator's GEO Data Report (Cloudflare Radar × 500+ site panel, July 21, 2026): Mistral is now the most extractive bot at 3,389:1, Anthropic improved 24× to 2,237:1, OpenAI collapsed to 217:1 as ChatGPT referrals crossed 1% of web traffic, and Google leads the majors at 4.6:1. Includes the month-by-month elasticity table, the 45.9% crawler 200-rate problem, and the exact "block training, allow search" robots.txt.

12 min read·Updated 2026-08-05

Robots.txt is the first gate between your content and AI search engines. If the crawlers that feed ChatGPT Search, Perplexity, Google AI Overviews, and Claude cannot fetch your pages, no amount of structured data or factual density will earn you a citation.

AI search engines use dedicated crawlers separate from their training crawlers. In 2026, there are thirteen major AI crawlers actively scanning the web, up from just five in 2024. The key strategic shift: website owners are increasingly adopting a "block training, allow search" posture — explicitly blocking extractive training bots that return no traffic while permitting search citation bots that drive referral visits. This guide covers every AI crawler you need to know, the latest 2026 blocking data, crawl-to-refer ratios, and the exact robots.txt configuration to deploy.

Why this matters in 2026: AI bot traffic now accounts for 33.21% of all web traffic (Q2 2026, up from 30.39% YoY). The 403 Forbidden rate on AI bots has more than doubled — from 3.63% in Q2 2025 to 8.56% in Q2 2026. Googlebot's share of AI bot traffic has been halved (57.2% → 27.49%) as the landscape fragments. ChatGPT Search handles 250–500 million weekly queries. The macro trend is unmistakable: the Imperva 2026 Bad Bot Report found that 53%+ of all web traffic is now automated (up from 51% in 2024), and Cloudflare data shows ~80% of AI bot crawling is for model training rather than live search. SE Ranking's 2026 study puts AI referral traffic at 0.32% of all web traffic — a 16× increase since 2024. Sites that explicitly block AI search crawlers are invisible in this channel regardless of content quality. HTTP Archive's 2025 crawl of 12.1M sites found 94.1% served a robots.txt, and by 2025 over 560,000 sites had explicit rules for ChatGPT, Claude, or Facebook AI agents (Paul Calvano / HTTP Archive) — yet a 2025 SERanking audit of 300,000 domains found that fewer than 4% had a correct OAI-SearchBot entry, silently blocking their own GEO potential.

The 13 AI crawlers you need to know in 2026

Each AI platform runs at least one crawler. Some run multiple: a training crawler that builds foundation models, a search crawler that powers real-time citations, and a user proxy that fetches pages on demand during conversations. The distinction matters — many sites block the training crawler and accidentally block the search crawler too. In Q2 2026, 87.51% of AI crawls are for training or mixed purposes, returning no direct value to publishers. Only 11.91% are for search or user action that can send referral traffic (Digital Applied, 2026).

User-agentOperatorPurposeBlocking affects citations?
OAI-SearchBotOpenAISearchGPT indexingYes
ChatGPT-UserOpenAIReal-time browsing in ChatGPTYes
GPTBotOpenAITraining + search indexingPartial
ClaudeBotAnthropicWeb content indexing for ClaudeYes
Claude-SearchBotAnthropicClaude search citationsYes
PerplexityBotPerplexityReal-time search indexingYes
Google-ExtendedGoogleGemini / Vertex AI trainingNo
Applebot-ExtendedAppleApple Intelligence trainingPartial
Meta-ExternalAgentMetaMeta AI features indexingPartial
BytespiderByteDanceTikTok / Doubao AI trainingNo
GoogleOtherGoogleMisc research / one-off tasksNo
CCBotCommon CrawlOpen web archive for trainingNo
MistralAI-UserMistralLe Chat training + retrievalNo

Sources: OpenAI platform documentation (2026), Anthropic crawler docs (2026), Google Search Central (2026), Perplexity documentation (2026), Digital Applied AI Crawler Statistics (2026), TechnologyChecker robots.txt analysis (2026).

AI crawler traffic share shifts (Q2 2025 → Q2 2026)

The AI crawler landscape has fragmented significantly over the past year. Googlebot's dominance has been cut in half as new AI crawlers scale up. This shift has major implications for robots.txt strategy — the old approach of "just allow Googlebot" is no longer sufficient for GEO.

AI BotQ2 2025 ShareQ2 2026 ShareChange
Googlebot57.20%27.49%−52%
ClaudeBot8.30%13.87%+67%
Meta-ExternalAgent6.27%12.70%+103%
GPTBot11.40%10.23%−10%
Bytespider2.37%7.91%+234%
Applebot0.92%7.36%+700%

Source: TechnologyChecker robots.txt analysis across Cloudflare's network (2026), Digital Applied AI Crawler & Bot Traffic Statistics (2026).

Crawl-to-refer ratios: the metric driving blocking decisions

The crawl-to-refer ratio — how many pages a crawler fetches for each referral visit it sends back — has become the key metric for surgical blocking decisions in 2026. The most rigorous measurement to date comes from SEOmator's GEO Data Report 2026 (July 21, 2026), which cross-checked Cloudflare Radar's global network against its own panel of 500+ sites. The two independent datasets agreed closely — Anthropic measured 2,237:1 on Cloudflare versus 2,363:1 on the SEOmator panel — confirming the signal is real rather than a sampling artifact.

The single most important finding: extraction ratios are elastic, not fixed. Every major extractor improved dramatically the moment it shipped a consumer product that links out. Anthropic fell from 56,969:1 in January 2026 to 2,363:1 by July — a 24× improvement in seven months. OpenAI collapsed from 1,264:1 to 179:1 over the same period as ChatGPT Search matured, and OpenAI referrals crossed 1% of all referral traffic (1.05% globally) ahead of schedule. The practical implication for robots.txt policy: a block decision made in January 2026 is almost certainly wrong by August. Re-audit quarterly.

OperatorCloudflare Radar (global)SEOmator panel (B2B)Blocking recommendation
Mistral (MistralAI-User)3,389:16,021:1Block with confidence
Anthropic (ClaudeBot)2,237:12,363:1Block training, allow Claude-SearchBot
Perplexity (PerplexityBot)225:1263:1Allow (but trending worse)
OpenAI (GPTBot + OAI-SearchBot)217:1179:1Allow OAI-SearchBot
Microsoft (Copilot / Bing)35:135:1Allow strategically
ByteDance (Bytespider)11:1Allow (referrals rising)
Google (Gemini / AI Overviews)4.6:14.0:1Allow
DuckDuckGo2.5:1Allow (most efficient of all)

Source: SEOmator, "GEO Data Report 2026: Which AI Crawlers & LLM Bots Take the Most and Give the Least?" (July 21, 2026) — Cloudflare Radar rolling 28-day window ending July 21, 2026, cross-checked against SEOmator Agent Analytics across 500+ sites (Jan 1 – Jul 21, 2026).

The elasticity trend: why quarterly re-audits matter

Tracking the same operators month by month shows how fast this metric moves. Every ratio below is from the SEOmator panel, so the numbers are directly comparable across columns:

OperatorJan 2026Mar 2026May 2026Jul 2026
Anthropic56,969:111,870:111,934:12,363:1
OpenAI1,264:11,416:11,225:1179:1
Perplexity116:1114:1175:1263:1
Mistral22:139:140:16,021:1
Google5.1:15.5:15.0:14.0:1

Source: SEOmator Agent Analytics panel, 500+ sites (July column covers Jul 1–21, 2026). Mistral's reversal reflects a large training crawl launched mid-2026 against a consumer surface that sends almost no clickable traffic back.

Two operators moved in opposite directions, and both moves are instructive. Anthropic and OpenAI improved by 24× and 7× respectively — not out of goodwill, but because they shipped cited-source surfaces that link out. Mistral went the other way, from a near-polite 22:1 in January to 6,021:1 in July, after launching a heavy training crawl behind a consumer product with no referral mechanism. Meta-ExternalAgent, now 6.35% of all crawler traffic, has never had a referral product at all.

"Extraction ratios are elastic, not fixed. The biggest extractors improved dramatically the moment they shipped consumer products that link out."
— Nick Sawinyh, SEOmator, "GEO Data Report 2026" (July 21, 2026)
"Google AI Overviews does not use a separate crawler. It reuses the regular Googlebot index. So if your site is in Google's index, you are already eligible for AI Overviews citations — provided your content matches the query."
— Google Search Central documentation, 2026
"More than half of all internet traffic is now bot-driven, and the fastest-growing slice is AI crawlers built to feed generative engines. Webmasters who treat robots.txt as a static afterthought are ceding citation visibility they will never get back."
Imperva, 2026 Bad Bot Report (bot traffic share trend)
"Authority decides whether you are in the candidate set; structure decides whether you are extracted. A clean robots.txt that lets the search crawlers in is the first line of structure — without it, no amount of E-E-A-T earns a citation."
Nico Digital, "AI Search Statistics 2026" (citation-outcome framework)

The minimum-viable robots.txt for GEO in 2026

The block below implements the "block training, allow search" posture — the consensus recommendation for 2026. It explicitly allows every AI search and citation crawler while blocking training-only crawlers. Paste it at the top of your robots.txt, above any User-agent: * block.

# ── AI citation crawlers: explicitly allow ──
User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Applebot-Extended
Allow: /

User-agent: Meta-ExternalAgent
Allow: /

# ── Training-only crawlers: block ──
User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

# ── Highest extraction, no referral product (July 2026 data) ──
User-agent: MistralAI-User
Disallow: /

# ── Default rules ──
User-agent: *
Disallow: /private/
Disallow: /admin/
Allow: /

Sitemap: https://geoaura.world/sitemap.xml

The three-layer AI visibility system

In 2026, robots.txt is no longer the only file governing AI interactions. A three-layer system has emerged as the industry standard for managing AI access:

FileRoleControls
robots.txtThe gatekeeperWhich crawlers can fetch which URLs
llms.txtThe tour guideAI-readable site context + content hierarchy
ai.txtThe policy documentAI usage rights, licensing, training opt-out policy

Robots.txt remains the access gatekeeper — if the crawler cannot reach your pages, nothing else matters. llms.txt provides a structured AI-readable map of your site's most important content. ai.txt declares your site's policy on AI training usage, content licensing, and citation permissions. Together, these three files form a complete AI visibility framework.

Why "block everything by default" backfires for GEO

A common SEO-school recommendation is to start robots.txt with User-agent: * / Disallow: / then selectively allow bots. This approach breaks GEO because most AI search crawlers fall back to the * rule when their specific block is missing or malformed. With AI crawler 403 forbidden rates at 8.56% (more than double the 3.63% rate in Q2 2025), accidental blocking is a growing and costly problem.

The TechnologyChecker 2026 analysis of Cloudflare's network found that the "block training, allow search" configuration is rapidly becoming the consensus standard. The key insight: AI crawlers fetch full page text paragraph-by-paragraph, looking for facts, claims, and quotable statements — they do not just index metadata and links like traditional search bots. A correctly configured robots.txt is the prerequisite for this extraction to happen at all.

Three rules for writing crawler blocks

  1. 1.
    Place specific agents before the wildcard.

    Robots.txt parsers match top-down. If User-agent: * comes first with a Disallow, the AI bot may inherit that block before reaching its own entry.

  2. 2.
    Each crawler needs its own block.

    You cannot list multiple crawlers under one User-agent directive. Each AI crawler requires its own User-agent line followed by Allow or Disallow.

  3. 3.
    Validate before deploying.

    Run your robots.txt through Google Search Console's robots.txt Tester and a third-party validator. A single misplaced character can silently block an entire crawler. Meta robots tags are not consistently respected by AI crawlers — robots.txt is the primary reliable gatekeeper.

The GPTBot caveat: a 2026 update

A critical nuance emerged in 2026: GPTBot serves dual purposes. It crawls pages for OpenAI model training and for search indexing. Blocking GPTBot while allowing OAI-SearchBot may reduce SearchGPT visibility because OpenAI may route a portion of search queries through GPTBot's index. The economics have shifted decisively in OpenAI's favour: its crawl-to-refer ratio fell from 1,264:1 in January 2026 to 179:1 by July, and ChatGPT referrals crossed 1% of all referral traffic — a threshold analysts had not expected until late 2026. OpenAI is now one of the few operators that returns meaningful traffic for what it takes, which weakens the case for blocking it at all. The conservative recommendation for maximum GEO impact: test with GPTBot blocked and monitor your ChatGPT Search citation rate for 30 days.

The silent killer: CDN bot-management rules

A correct robots.txt is necessary but not sufficient. Nico Digital's July 2026 audit flags a growing failure mode: CDN and WAF bot-management rules silently block AI crawlers by user-agent, excluding pages from OpenAI's index even when robots.txt explicitly allows OAI-SearchBot and GPTBot. The crawler passes the robots.txt gate, then gets a 403 at the edge — and your content never enters the retrieval pool.

Two layers to check: (1) robots.txt allows the search crawler, and (2) your CDN/WAF allowlist permits that same user-agent. With AI crawler 403 rates at 8.56% (more than double Q2 2025), a surprising share of "missing citations" trace to edge-layer blocks, not content quality. The scale of this failure is larger than most teams assume: SEOmator's July 2026 analysis found only 45.9% of all crawler requests return a 200, and for AI crawlers specifically just ~46% get a 200 OK while more than a quarter are blocked or throttled outright. Audit both layers before blaming GEO strategy.

The underrated AEO target: Bing Copilot + IndexNow

Because ChatGPT Search retrieves from the Bing index, Bing Copilot is the most underrated AI citation surface of 2026. Bing's lower SEO competition and native IndexNow integration — which pushes new and updated URLs to Bing within hours — make it the fastest way to get fresh content into the pool that powers both Bing Copilot and ChatGPT Search. Submitting via IndexNow after every update (as this site does on publish) shortens the citation lag from weeks to days.

Verifying crawler access

After deploying, verify that each AI search crawler can actually fetch your pages. Three checks cover 95% of issues:

  • Server logs — grep your access logs for the user-agent strings above. If you see fetches, the bot reached you. Bytespider is known for aggressive crawl rates — monitor your log volume.
  • robots.txt Tester — Google Search Console lets you test any user-agent against your live robots.txt. Test "OAI-SearchBot", "ChatGPT-User", "PerplexityBot", and "ClaudeBot" explicitly.
  • Direct citation check — ask ChatGPT Search and Perplexity a question your site should answer. If neither cites you after 4 weeks of correct robots.txt, the problem is content quality, not access.

Frequently asked questions

Which AI crawler user-agents should I allow in robots.txt?

At minimum allow OAI-SearchBot (ChatGPT Search), ChatGPT-User (ChatGPT real-time browsing), PerplexityBot, ClaudeBot (Claude citations), Applebot-Extended (Apple Intelligence), and Meta-ExternalAgent (Meta AI). GPTBot and Google-Extended are training crawlers and optional if you only want search citation.

Does Google AI Overviews use a separate crawler?

No. Google AI Overviews reuses the regular Googlebot index, so you do not need a new user-agent entry. Google-Extended is only used to opt Gemini training data in or out.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot crawls pages for OpenAI model training and search indexing. OAI-SearchBot crawls pages specifically to power ChatGPT Search inline citations. Blocking GPTBot while allowing OAI-SearchBot may reduce SearchGPT visibility because OpenAI may use GPTBot for underlying search indexing.

What is the crawl-to-refer ratio and why does it matter?

It measures how many pages a crawler fetches for each referral visit it sends back. In July 2026, Mistral was the most extractive at 3,389:1, ahead of Anthropic (2,237:1), Perplexity (225:1) and OpenAI (217:1); Google was the most efficient major at 4.6:1 and DuckDuckGo led all operators at 2.5:1. Sites use this metric to decide which crawlers to block or allow, adopting a "block training, allow search" posture. Because the ratios move fast — Anthropic improved 24× in seven months — re-audit quarterly rather than setting the policy once.

What is the three-layer AI visibility system?

Robots.txt controls crawler access (the gatekeeper). llms.txt provides AI-readable site context (the tour guide). ai.txt declares AI usage and licensing policies (the policy document). All three work together as the recommended framework for managing AI access in 2026.

References: SEOmator — "GEO Data Report 2026: Which AI Crawlers & LLM Bots Take the Most and Give the Least?" (Nick Sawinyh, July 21, 2026): crawl-to-refer ratios cross-checked between Cloudflare Radar's global network (rolling 28-day window ending Jul 21, 2026) and SEOmator Agent Analytics across 500+ sites (Jan 1 – Jul 21, 2026); robots.txt parse of 4,257 top domains (Jul 20, 2026); AI-readiness scans of ~106,000 (Cloudflare) and ~109,000 (SEOmator) domains. · SEOmator — "Crawl Waste Report 2026: Only 45.9% of Crawler Requests Return a 200" (July 19, 2026) and "AI Bot Traffic by Country: Where AI Crawlers Are Most Aggressive" (July 24, 2026). · OpenAI platform docs — GPTBot, ChatGPT-User & OAI-SearchBot (2026). · Anthropic documentation — ClaudeBot & Claude-SearchBot (2026). · Google Search Central — Google-Extended (2026). · Perplexity documentation — PerplexityBot (2026). · Meta crawler documentation — Meta-ExternalAgent (2026). · TechnologyChecker — "Robots.txt & AI Crawlers Blocking Report" (Cloudflare network analysis, 2026). · Digital Applied — "AI Crawler & Bot Traffic Statistics 2026" (Cloudflare Radar data, Imperva Bad Bot Report). · Imperva — 2026 Bad Bot Report (bot traffic share: 53%+ automated). · Cloudflare — AI crawler traffic & crawl-to-refer data (2025–2026). · SE Ranking — AI Traffic Research Study 2026 (16× referral growth). · Paul Calvano / HTTP Archive — robots.txt & AI-agent adoption across 12.1M sites (2025). · Nico Digital — "AI Search Statistics 2026" (citation-outcome framework). · Similarweb 2026 AI Search Report. · Presenc AI — AI Search Engine Market Share 2026. · SERanking robots.txt audit of 300,000 domains (2025). · Aggarwal et al., "GEO: Generative Engine Optimization," arXiv:2311.09735, KDD 2024.

Want to check your site's GEO readiness?

Run the 27-point GEO audit