AI Crawler Robots.txt Guide: Every Bot You Need to Allow (2026 Update)
13 AI crawlers now scan the web — OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Meta-ExternalAgent, MistralAI-User, and more. Rebuilt August 2026 on SEOmator's GEO Data Report (Cloudflare Radar × 500+ site panel, July 21, 2026): Mistral is now the most extractive bot at 3,389:1, Anthropic improved 24× to 2,237:1, OpenAI collapsed to 217:1 as ChatGPT referrals crossed 1% of web traffic, and Google leads the majors at 4.6:1. Includes the month-by-month elasticity table, the 45.9% crawler 200-rate problem, and the exact "block training, allow search" robots.txt.
Robots.txt is the first gate between your content and AI search engines. If the crawlers that feed ChatGPT Search, Perplexity, Google AI Overviews, and Claude cannot fetch your pages, no amount of structured data or factual density will earn you a citation.
AI search engines use dedicated crawlers separate from their training crawlers. In 2026, there are thirteen major AI crawlers actively scanning the web, up from just five in 2024. The key strategic shift: website owners are increasingly adopting a "block training, allow search" posture — explicitly blocking extractive training bots that return no traffic while permitting search citation bots that drive referral visits. This guide covers every AI crawler you need to know, the latest 2026 blocking data, crawl-to-refer ratios, and the exact robots.txt configuration to deploy.
Why this matters in 2026: AI bot traffic now accounts for 33.21% of all web traffic (Q2 2026, up from 30.39% YoY). The 403 Forbidden rate on AI bots has more than doubled — from 3.63% in Q2 2025 to 8.56% in Q2 2026. Googlebot's share of AI bot traffic has been halved (57.2% → 27.49%) as the landscape fragments. ChatGPT Search handles 250–500 million weekly queries. The macro trend is unmistakable: the Imperva 2026 Bad Bot Report found that 53%+ of all web traffic is now automated (up from 51% in 2024), and Cloudflare data shows ~80% of AI bot crawling is for model training rather than live search. SE Ranking's 2026 study puts AI referral traffic at 0.32% of all web traffic — a 16× increase since 2024. Sites that explicitly block AI search crawlers are invisible in this channel regardless of content quality. HTTP Archive's 2025 crawl of 12.1M sites found 94.1% served a robots.txt, and by 2025 over 560,000 sites had explicit rules for ChatGPT, Claude, or Facebook AI agents (Paul Calvano / HTTP Archive) — yet a 2025 SERanking audit of 300,000 domains found that fewer than 4% had a correct OAI-SearchBot entry, silently blocking their own GEO potential.
The 13 AI crawlers you need to know in 2026
Each AI platform runs at least one crawler. Some run multiple: a training crawler that builds foundation models, a search crawler that powers real-time citations, and a user proxy that fetches pages on demand during conversations. The distinction matters — many sites block the training crawler and accidentally block the search crawler too. In Q2 2026, 87.51% of AI crawls are for training or mixed purposes, returning no direct value to publishers. Only 11.91% are for search or user action that can send referral traffic (Digital Applied, 2026).
| User-agent | Operator | Purpose | Blocking affects citations? |
|---|---|---|---|
| OAI-SearchBot | OpenAI | SearchGPT indexing | Yes |
| ChatGPT-User | OpenAI | Real-time browsing in ChatGPT | Yes |
| GPTBot | OpenAI | Training + search indexing | Partial |
| ClaudeBot | Anthropic | Web content indexing for Claude | Yes |
| Claude-SearchBot | Anthropic | Claude search citations | Yes |
| PerplexityBot | Perplexity | Real-time search indexing | Yes |
| Google-Extended | Gemini / Vertex AI training | No | |
| Applebot-Extended | Apple | Apple Intelligence training | Partial |
| Meta-ExternalAgent | Meta | Meta AI features indexing | Partial |
| Bytespider | ByteDance | TikTok / Doubao AI training | No |
| GoogleOther | Misc research / one-off tasks | No | |
| CCBot | Common Crawl | Open web archive for training | No |
| MistralAI-User | Mistral | Le Chat training + retrieval | No |
Sources: OpenAI platform documentation (2026), Anthropic crawler docs (2026), Google Search Central (2026), Perplexity documentation (2026), Digital Applied AI Crawler Statistics (2026), TechnologyChecker robots.txt analysis (2026).
AI crawler traffic share shifts (Q2 2025 → Q2 2026)
The AI crawler landscape has fragmented significantly over the past year. Googlebot's dominance has been cut in half as new AI crawlers scale up. This shift has major implications for robots.txt strategy — the old approach of "just allow Googlebot" is no longer sufficient for GEO.
| AI Bot | Q2 2025 Share | Q2 2026 Share | Change |
|---|---|---|---|
| Googlebot | 57.20% | 27.49% | −52% |
| ClaudeBot | 8.30% | 13.87% | +67% |
| Meta-ExternalAgent | 6.27% | 12.70% | +103% |
| GPTBot | 11.40% | 10.23% | −10% |
| Bytespider | 2.37% | 7.91% | +234% |
| Applebot | 0.92% | 7.36% | +700% |
Source: TechnologyChecker robots.txt analysis across Cloudflare's network (2026), Digital Applied AI Crawler & Bot Traffic Statistics (2026).
Crawl-to-refer ratios: the metric driving blocking decisions
The crawl-to-refer ratio — how many pages a crawler fetches for each referral visit it sends back — has become the key metric for surgical blocking decisions in 2026. The most rigorous measurement to date comes from SEOmator's GEO Data Report 2026 (July 21, 2026), which cross-checked Cloudflare Radar's global network against its own panel of 500+ sites. The two independent datasets agreed closely — Anthropic measured 2,237:1 on Cloudflare versus 2,363:1 on the SEOmator panel — confirming the signal is real rather than a sampling artifact.
The single most important finding: extraction ratios are elastic, not fixed. Every major extractor improved dramatically the moment it shipped a consumer product that links out. Anthropic fell from 56,969:1 in January 2026 to 2,363:1 by July — a 24× improvement in seven months. OpenAI collapsed from 1,264:1 to 179:1 over the same period as ChatGPT Search matured, and OpenAI referrals crossed 1% of all referral traffic (1.05% globally) ahead of schedule. The practical implication for robots.txt policy: a block decision made in January 2026 is almost certainly wrong by August. Re-audit quarterly.
| Operator | Cloudflare Radar (global) | SEOmator panel (B2B) | Blocking recommendation |
|---|---|---|---|
| Mistral (MistralAI-User) | 3,389:1 | 6,021:1 | Block with confidence |
| Anthropic (ClaudeBot) | 2,237:1 | 2,363:1 | Block training, allow Claude-SearchBot |
| Perplexity (PerplexityBot) | 225:1 | 263:1 | Allow (but trending worse) |
| OpenAI (GPTBot + OAI-SearchBot) | 217:1 | 179:1 | Allow OAI-SearchBot |
| Microsoft (Copilot / Bing) | 35:1 | 35:1 | Allow strategically |
| ByteDance (Bytespider) | 11:1 | — | Allow (referrals rising) |
| Google (Gemini / AI Overviews) | 4.6:1 | 4.0:1 | Allow |
| DuckDuckGo | 2.5:1 | — | Allow (most efficient of all) |
Source: SEOmator, "GEO Data Report 2026: Which AI Crawlers & LLM Bots Take the Most and Give the Least?" (July 21, 2026) — Cloudflare Radar rolling 28-day window ending July 21, 2026, cross-checked against SEOmator Agent Analytics across 500+ sites (Jan 1 – Jul 21, 2026).
The elasticity trend: why quarterly re-audits matter
Tracking the same operators month by month shows how fast this metric moves. Every ratio below is from the SEOmator panel, so the numbers are directly comparable across columns:
| Operator | Jan 2026 | Mar 2026 | May 2026 | Jul 2026 |
|---|---|---|---|---|
| Anthropic | 56,969:1 | 11,870:1 | 11,934:1 | 2,363:1 |
| OpenAI | 1,264:1 | 1,416:1 | 1,225:1 | 179:1 |
| Perplexity | 116:1 | 114:1 | 175:1 | 263:1 |
| Mistral | 22:1 | 39:1 | 40:1 | 6,021:1 |
| 5.1:1 | 5.5:1 | 5.0:1 | 4.0:1 |
Source: SEOmator Agent Analytics panel, 500+ sites (July column covers Jul 1–21, 2026). Mistral's reversal reflects a large training crawl launched mid-2026 against a consumer surface that sends almost no clickable traffic back.
Two operators moved in opposite directions, and both moves are instructive. Anthropic and OpenAI improved by 24× and 7× respectively — not out of goodwill, but because they shipped cited-source surfaces that link out. Mistral went the other way, from a near-polite 22:1 in January to 6,021:1 in July, after launching a heavy training crawl behind a consumer product with no referral mechanism. Meta-ExternalAgent, now 6.35% of all crawler traffic, has never had a referral product at all.
"Extraction ratios are elastic, not fixed. The biggest extractors improved dramatically the moment they shipped consumer products that link out."
"Google AI Overviews does not use a separate crawler. It reuses the regular Googlebot index. So if your site is in Google's index, you are already eligible for AI Overviews citations — provided your content matches the query."
"More than half of all internet traffic is now bot-driven, and the fastest-growing slice is AI crawlers built to feed generative engines. Webmasters who treat robots.txt as a static afterthought are ceding citation visibility they will never get back."
"Authority decides whether you are in the candidate set; structure decides whether you are extracted. A clean robots.txt that lets the search crawlers in is the first line of structure — without it, no amount of E-E-A-T earns a citation."
The minimum-viable robots.txt for GEO in 2026
The block below implements the "block training, allow search" posture — the consensus recommendation for 2026. It explicitly allows every AI search and citation crawler while blocking training-only crawlers. Paste it at the top of your robots.txt, above any User-agent: * block.
# ── AI citation crawlers: explicitly allow ── User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: PerplexityBot Allow: / User-agent: ClaudeBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Applebot-Extended Allow: / User-agent: Meta-ExternalAgent Allow: / # ── Training-only crawlers: block ── User-agent: GPTBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / # ── Highest extraction, no referral product (July 2026 data) ── User-agent: MistralAI-User Disallow: / # ── Default rules ── User-agent: * Disallow: /private/ Disallow: /admin/ Allow: / Sitemap: https://geoaura.world/sitemap.xml
The three-layer AI visibility system
In 2026, robots.txt is no longer the only file governing AI interactions. A three-layer system has emerged as the industry standard for managing AI access:
| File | Role | Controls |
|---|---|---|
| robots.txt | The gatekeeper | Which crawlers can fetch which URLs |
| llms.txt | The tour guide | AI-readable site context + content hierarchy |
| ai.txt | The policy document | AI usage rights, licensing, training opt-out policy |
Robots.txt remains the access gatekeeper — if the crawler cannot reach your pages, nothing else matters. llms.txt provides a structured AI-readable map of your site's most important content. ai.txt declares your site's policy on AI training usage, content licensing, and citation permissions. Together, these three files form a complete AI visibility framework.
Why "block everything by default" backfires for GEO
A common SEO-school recommendation is to start robots.txt with User-agent: * / Disallow: / then selectively allow bots. This approach breaks GEO because most AI search crawlers fall back to the * rule when their specific block is missing or malformed. With AI crawler 403 forbidden rates at 8.56% (more than double the 3.63% rate in Q2 2025), accidental blocking is a growing and costly problem.
The TechnologyChecker 2026 analysis of Cloudflare's network found that the "block training, allow search" configuration is rapidly becoming the consensus standard. The key insight: AI crawlers fetch full page text paragraph-by-paragraph, looking for facts, claims, and quotable statements — they do not just index metadata and links like traditional search bots. A correctly configured robots.txt is the prerequisite for this extraction to happen at all.
Three rules for writing crawler blocks
- 1.Place specific agents before the wildcard.
Robots.txt parsers match top-down. If
User-agent: *comes first with a Disallow, the AI bot may inherit that block before reaching its own entry. - 2.Each crawler needs its own block.
You cannot list multiple crawlers under one User-agent directive. Each AI crawler requires its own
User-agentline followed byAlloworDisallow. - 3.Validate before deploying.
Run your robots.txt through Google Search Console's robots.txt Tester and a third-party validator. A single misplaced character can silently block an entire crawler. Meta robots tags are not consistently respected by AI crawlers — robots.txt is the primary reliable gatekeeper.
The GPTBot caveat: a 2026 update
A critical nuance emerged in 2026: GPTBot serves dual purposes. It crawls pages for OpenAI model training and for search indexing. Blocking GPTBot while allowing OAI-SearchBot may reduce SearchGPT visibility because OpenAI may route a portion of search queries through GPTBot's index. The economics have shifted decisively in OpenAI's favour: its crawl-to-refer ratio fell from 1,264:1 in January 2026 to 179:1 by July, and ChatGPT referrals crossed 1% of all referral traffic — a threshold analysts had not expected until late 2026. OpenAI is now one of the few operators that returns meaningful traffic for what it takes, which weakens the case for blocking it at all. The conservative recommendation for maximum GEO impact: test with GPTBot blocked and monitor your ChatGPT Search citation rate for 30 days.
The silent killer: CDN bot-management rules
A correct robots.txt is necessary but not sufficient. Nico Digital's July 2026 audit flags a growing failure mode: CDN and WAF bot-management rules silently block AI crawlers by user-agent, excluding pages from OpenAI's index even when robots.txt explicitly allows OAI-SearchBot and GPTBot. The crawler passes the robots.txt gate, then gets a 403 at the edge — and your content never enters the retrieval pool.
Two layers to check: (1) robots.txt allows the search crawler, and (2) your CDN/WAF allowlist permits that same user-agent. With AI crawler 403 rates at 8.56% (more than double Q2 2025), a surprising share of "missing citations" trace to edge-layer blocks, not content quality. The scale of this failure is larger than most teams assume: SEOmator's July 2026 analysis found only 45.9% of all crawler requests return a 200, and for AI crawlers specifically just ~46% get a 200 OK while more than a quarter are blocked or throttled outright. Audit both layers before blaming GEO strategy.
The underrated AEO target: Bing Copilot + IndexNow
Because ChatGPT Search retrieves from the Bing index, Bing Copilot is the most underrated AI citation surface of 2026. Bing's lower SEO competition and native IndexNow integration — which pushes new and updated URLs to Bing within hours — make it the fastest way to get fresh content into the pool that powers both Bing Copilot and ChatGPT Search. Submitting via IndexNow after every update (as this site does on publish) shortens the citation lag from weeks to days.
Verifying crawler access
After deploying, verify that each AI search crawler can actually fetch your pages. Three checks cover 95% of issues:
- ▸ Server logs — grep your access logs for the user-agent strings above. If you see fetches, the bot reached you. Bytespider is known for aggressive crawl rates — monitor your log volume.
- ▸ robots.txt Tester — Google Search Console lets you test any user-agent against your live robots.txt. Test "OAI-SearchBot", "ChatGPT-User", "PerplexityBot", and "ClaudeBot" explicitly.
- ▸ Direct citation check — ask ChatGPT Search and Perplexity a question your site should answer. If neither cites you after 4 weeks of correct robots.txt, the problem is content quality, not access.
Frequently asked questions
Which AI crawler user-agents should I allow in robots.txt?
At minimum allow OAI-SearchBot (ChatGPT Search), ChatGPT-User (ChatGPT real-time browsing), PerplexityBot, ClaudeBot (Claude citations), Applebot-Extended (Apple Intelligence), and Meta-ExternalAgent (Meta AI). GPTBot and Google-Extended are training crawlers and optional if you only want search citation.
Does Google AI Overviews use a separate crawler?
No. Google AI Overviews reuses the regular Googlebot index, so you do not need a new user-agent entry. Google-Extended is only used to opt Gemini training data in or out.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot crawls pages for OpenAI model training and search indexing. OAI-SearchBot crawls pages specifically to power ChatGPT Search inline citations. Blocking GPTBot while allowing OAI-SearchBot may reduce SearchGPT visibility because OpenAI may use GPTBot for underlying search indexing.
What is the crawl-to-refer ratio and why does it matter?
It measures how many pages a crawler fetches for each referral visit it sends back. In July 2026, Mistral was the most extractive at 3,389:1, ahead of Anthropic (2,237:1), Perplexity (225:1) and OpenAI (217:1); Google was the most efficient major at 4.6:1 and DuckDuckGo led all operators at 2.5:1. Sites use this metric to decide which crawlers to block or allow, adopting a "block training, allow search" posture. Because the ratios move fast — Anthropic improved 24× in seven months — re-audit quarterly rather than setting the policy once.
What is the three-layer AI visibility system?
Robots.txt controls crawler access (the gatekeeper). llms.txt provides AI-readable site context (the tour guide). ai.txt declares AI usage and licensing policies (the policy document). All three work together as the recommended framework for managing AI access in 2026.
References: SEOmator — "GEO Data Report 2026: Which AI Crawlers & LLM Bots Take the Most and Give the Least?" (Nick Sawinyh, July 21, 2026): crawl-to-refer ratios cross-checked between Cloudflare Radar's global network (rolling 28-day window ending Jul 21, 2026) and SEOmator Agent Analytics across 500+ sites (Jan 1 – Jul 21, 2026); robots.txt parse of 4,257 top domains (Jul 20, 2026); AI-readiness scans of ~106,000 (Cloudflare) and ~109,000 (SEOmator) domains. · SEOmator — "Crawl Waste Report 2026: Only 45.9% of Crawler Requests Return a 200" (July 19, 2026) and "AI Bot Traffic by Country: Where AI Crawlers Are Most Aggressive" (July 24, 2026). · OpenAI platform docs — GPTBot, ChatGPT-User & OAI-SearchBot (2026). · Anthropic documentation — ClaudeBot & Claude-SearchBot (2026). · Google Search Central — Google-Extended (2026). · Perplexity documentation — PerplexityBot (2026). · Meta crawler documentation — Meta-ExternalAgent (2026). · TechnologyChecker — "Robots.txt & AI Crawlers Blocking Report" (Cloudflare network analysis, 2026). · Digital Applied — "AI Crawler & Bot Traffic Statistics 2026" (Cloudflare Radar data, Imperva Bad Bot Report). · Imperva — 2026 Bad Bot Report (bot traffic share: 53%+ automated). · Cloudflare — AI crawler traffic & crawl-to-refer data (2025–2026). · SE Ranking — AI Traffic Research Study 2026 (16× referral growth). · Paul Calvano / HTTP Archive — robots.txt & AI-agent adoption across 12.1M sites (2025). · Nico Digital — "AI Search Statistics 2026" (citation-outcome framework). · Similarweb 2026 AI Search Report. · Presenc AI — AI Search Engine Market Share 2026. · SERanking robots.txt audit of 300,000 domains (2025). · Aggarwal et al., "GEO: Generative Engine Optimization," arXiv:2311.09735, KDD 2024.
Want to check your site's GEO readiness?
Run the 27-point GEO auditRelated articles
AI Search Engines Compared 2026: Perplexity vs ChatGPT vs Gemini vs Claude vs Grok
Comprehensive comparison of the five major AI search engines as of July 2026. Perplexity wins on citation honesty, Claude on synthesis depth, Gemini on real-time freshness (+9× growth, ~750M MAU), ChatGPT on multi-source reasoning (900M WAU), and Grok on speed. AI search now processes 3.5B+ queries/week (Axis Intelligence); AI Overviews cover up to 48% of queries. With fresh market data and use-case recommendations.
ChatGPT Search Citation 2026: How It Cites Sources & How to Get Cited
ChatGPT Search uses OAI-SearchBot and inline citations, allocating only 3–8 source slots per answer. With 900M+ weekly active users, 76.85% of AI referral traffic, and Gartner predicting 25% of desktop search shifting to AI agents by 2026, this 2026 guide covers GPT-5.5, the May 2026 link update (+157.7% referral boost), Deep Research, and the 5-step ChatGPT search optimization checklist to get your content cited.
Perplexity AI Citation 2026: Mechanism & Optimization Guide
Perplexity uses 5-15 numbered references per answer — the most citation-dense AI engine. Updated with 2026 data: 100M+ MAU, 7.73% market share, Comet browser, Deep Research, the end-2026 ad wind-down pivot, and the complete optimization framework for Perplexity citations.