Last updated: 28 September 2026
AI Crawlers List 2026: Every Bot UK Businesses Need to Know
An AI crawlers list catalogues the automated bots — such as GPTBot, ClaudeBot, Google-Extended and PerplexityBot — that AI companies use to scrape UK websites for model training or real-time search answers. In 2026, bots account for 57.5% of HTML web traffic, overtaking human visitors for the first time, according to Cloudflare Radar via Digital Applied.
Key Takeaways
- GPTBot is the most-blocked AI crawler globally, appearing in 5.52% of robots.txt DISALLOW rules in Q1 2026, ahead of CCBot, ClaudeBot, Google-Extended and Bytespider, Aether AI notes, citing Technology Checker via Digital Applied.
- Aether AI highlights that bots now outnumber humans on the open web: Cloudflare Radar recorded bots at 57.5% of HTML traffic versus 42.5% human traffic in mid-2026, per Cloudflare Radar.
- Aether AI reports that Anthropic's ClaudeBot crawls roughly 2,237 pages for every visitor it refers, compared with about 217 pages per referral for OpenAI's GPTBot and 4.6 for Google, according to Cloudflare Radar via Heybuffy.
- Cloudflare made blocking AI crawlers the default setting for new domains from 1 July 2026, and customers blocked 416 billion AI scraping requests in the following five months, per Cloudflare Press Release and Digital Applied.
- GPTBot's share of verified bot traffic grew from 4.7% in July 2026 to 11.7% in July 2026, according to Cloudflare via Digital Applied.
Which companies operate AI crawlers and what are their user-agent names?
An AI crawler is an automated bot, identified by a specific user-agent string in server logs, that AI companies deploy to fetch web pages either for model training or for answering live user queries. Every major AI lab now runs at least one named crawler, and most UK site owners will see several of these in their logs weekly.
The best-known operators and their user-agent tokens include:
- OpenAI — GPTBot (training), OAI-SearchBot (search indexing), ChatGPT-User (live user requests inside ChatGPT)
- Anthropic — ClaudeBot and Claude-Web
- Google — Google-Extended (a control token separate from Googlebot, governing Gemini and AI Overviews training use)
- Perplexity — PerplexityBot, plus an undeclared crawler flagged by researchers for ignoring robots.txt
- ByteDance — Bytespider
- Common Crawl — CCBot, whose data feeds many downstream AI models
- Apple — Applebot-Extended
- Microsoft — Bingbot (partly used for Copilot) and BingPreview
The OpenAI Platform documentation and the community-maintained ai.robots.txt project on GitHub are the two most reliable places to confirm current tokens, since new crawlers appear regularly.
What is the purpose of each AI crawler?
AI crawler purpose falls into two broad categories: training data collection and real-time retrieval for live answers. GPTBot and ClaudeBot exist primarily to harvest text for training future model versions, while OAI-SearchBot, ChatGPT-User and PerplexityBot fetch pages on demand to answer a specific user query in the moment.
This distinction matters because the two purposes carry very different traffic patterns. Training crawlers sweep large volumes of pages in bulk, often with no near-term benefit to the site owner. Retrieval crawlers fetch a page because a user just asked something related to it, which can translate into a citation and, occasionally, referral traffic.
Cloudflare's crawl-to-refer data illustrates the gap starkly. Anthropic's ClaudeBot crawled roughly 2,237 pages for every visitor it referred, OpenAI's GPTBot around 217, and Google about 4.6, in the 28-day window ending 21 July 2026, according to Cloudflare Radar via Heybuffy. Compared with traditional Google Search, OpenAI's tools drive 750 times less referral traffic, while Anthropic's bots drive 30,000 times less, per Cloudflare internal metrics via TechNewsHub.
How can I identify AI crawler visits in my UK website's server logs?
Identifying AI crawler visits means filtering server logs or analytics for known user-agent strings such as GPTBot, ClaudeBot or PerplexityBot, then cross-checking the requesting IP address against each operator's published range. Most standard analytics tools such as Google Analytics 4 filter out bot traffic by default, so log-level inspection is usually necessary.
The practical steps for a UK site owner are straightforward:
- Export raw access logs from your hosting provider, CDN or WAF (web application firewall).
- Search the
user-agentfield for known tokens (GPTBot, ClaudeBot, Google-Extended, PerplexityBot, Bytespider, CCBot). - Verify the source IP against the operator's published range — OpenAI, Google and Anthropic all publish IP blocks or verification methods, because user-agent strings alone can be spoofed.
- Note request frequency and depth — a crawler hitting thousands of URLs in a short burst behaves very differently from a normal visitor.
Fact-check tip: never trust the user-agent string alone. Bytespider and some undeclared crawlers are known to ignore robots.txt entirely, so IP verification at the firewall level is the only reliable confirmation for high-value sites.
How do I block or allow specific AI crawlers using robots.txt?
Robots.txt is a plain-text file, placed at the root of a domain, that instructs compliant crawlers which parts of a site they may or may not access using User-agent and Disallow directives. Compliant AI crawlers such as GPTBot and ClaudeBot respect this file; some, including certain Perplexity and Bytespider requests, have been reported to ignore it.
A typical blocking rule looks like this:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
To allow a crawler instead, simply omit its block or explicitly permit it with Disallow: left blank. Fourteen per cent of top domains now use robots.txt rules specifically to manage AI and search crawlers, according to the Cloudflare Blog, which also found AI and search crawler traffic rose 18% (48% including new Cloudflare customers) between May 2026 and May 2026, with GPTBot growing 305% and Googlebot 96% in the same period.
Robots.txt is advisory, not enforceable. For crawlers known to ignore it, a WAF rule blocking the verified IP range is the only dependable control.
Does blocking AI crawlers affect visibility in ChatGPT, Perplexity or AI Overviews?
Blocking an AI crawler prevents that operator's assistant from citing your content in live answers, because tools like ChatGPT, Perplexity and Google's AI Overviews can only surface pages they are permitted to fetch or that exist in their training data. Disallowing GPTBot, for example, removes a UK business from OpenAI's retrieval pool for that domain going forward.
This creates a genuine trade-off. Cloudflare made blocking AI crawlers the default setting for all new domains from 1 July 2026, and by August 2026 more than 2.5 million sites had chosen to fully disallow AI training, according to the Cloudflare Press Release and Digital Applied. Cloudflare co-founder and CEO Matthew Prince has argued: "Instead of being a fair trade, the web is being strip-mined by AI crawlers. Content creators see almost no traffic and therefore almost no value," per Cloudflare.
"Blocking GPTBot to protect your content is like delisting from Google to protect your brochure. Defensible for a paywalled publisher; self-harm for anyone who sells something. If assistants can't read you, they recommend whoever they can read." — Lauren Dawkins, Head of Content, Aether AI
Across 60 client sites tracked between January and June 2026, Aether AI recorded that domains blocking OAI-SearchBot or PerplexityBot saw zero citations from the corresponding engine over the six-month window, consistent with the broader industry pattern that retrieval crawlers must reach a page to cite it.
Training crawlers vs retrieval crawlers: the decision
| Crawler type | Example bots | Blocks training? | Blocks live citations? | Typical decision |
|---|---|---|---|---|
| Training-only | GPTBot, CCBot, Bytespider | Yes | No (separate token) | Many publishers block |
| Retrieval/search | OAI-SearchBot, ChatGPT-User, PerplexityBot | N/A | Yes, if blocked | Most commercial sites allow |
| Google-Extended | Google-Extended | Yes (Gemini/AI training) | Does not affect classic Search ranking | Depends on content sensitivity |
UK GDPR and copyright implications of AI crawlers scraping content
UK GDPR applies to AI crawlers whenever the scraped pages contain personal data, because the Information Commissioner's Office (ICO) treats web scraping of identifiable information as processing under the UK GDPR and the Data Protection Act 2018. A site listing named staff, client testimonials with full names, or user-generated comments may be transferring personal data to an AI company's training pipeline without a clear lawful basis.
Separately, UK copyright law under the Copyright, Designs and Patents Act 1988 protects original written content, and copying substantial portions via crawling without a licence or a valid exception may constitute infringement. There is no general UK text-and-data-mining exception for commercial AI training equivalent to some EU provisions, which is why several publishers have pursued licensing deals with AI companies rather than relying solely on robots.txt.
Practical steps for UK businesses:
- Audit pages containing personal data (staff bios, testimonials, forms) before deciding on crawler access.
- Review your privacy notice to reflect third-party scraping risk.
- Consider a licensing or pay-per-crawl arrangement where available, rather than a blanket block, if content has commercial value to AI platforms.
Your AI crawlers list checklist
- Pull a current AI crawler user-agent list from ai.robots.txt on GitHub or the OpenAI bots documentation — this list changes monthly.
- Export server logs and filter for GPTBot, ClaudeBot, Google-Extended, PerplexityBot, Bytespider and CCBot user-agents.
- Verify suspected crawler IPs against each operator's published ranges before trusting the user-agent string alone.
- Decide separately on training crawlers (GPTBot, CCBot) versus retrieval crawlers (OAI-SearchBot, PerplexityBot, ChatGPT-User).
- Update robots.txt with explicit
User-agent/Disallowblocks for anything you want to exclude. - Add WAF-level IP blocking for non-compliant crawlers such as Bytespider or undeclared Perplexity requests.
- Audit pages with personal data for UK GDPR exposure before allowing broad scraping.
- Re-check your crawler list and robots.txt quarterly, since new bots and tokens appear regularly.
FAQ
What is an AI crawlers list and why does it matter for GEO?
An AI crawlers list is a catalogue of the user-agent names AI companies use to identify their bots, such as GPTBot, ClaudeBot and PerplexityBot. It matters for GEO (generative engine optimisation) because knowing which crawlers are active lets a business decide which ones to allow for AI search citations and which to block for training.
How do I block GPTBot specifically using robots.txt?
Add User-agent: GPTBot followed by Disallow: / to your robots.txt file at the domain root. This instructs OpenAI's training crawler not to fetch any pages, though it does not block OAI-SearchBot or ChatGPT-User, which use separate tokens.
Does blocking Google-Extended affect my normal Google Search rankings?
No, blocking Google-Extended does not affect classic Google Search ranking, because it is a separate control token from Googlebot used specifically for Gemini and AI Overviews training. Googlebot continues indexing your site normally unless blocked separately.
How can I tell if an AI crawler visit is genuine or spoofed?
Check the requesting IP address against the operator's officially published IP range, since user-agent strings alone can be faked by other bots or scrapers. Legitimate operators such as OpenAI, Google and Anthropic publish verification methods for this purpose.
What happens if I block all AI crawlers on my website?
Blocking all AI crawlers removes your site from training datasets and, for retrieval crawlers, prevents citations in tools like ChatGPT, Perplexity and Google AI Overviews. Given that AI referral traffic is already far smaller than traditional search referral traffic, the visibility trade-off depends heavily on your business model.
How often does the AI crawler user-agent list change?
The list changes frequently, with new tokens and operators appearing every few months as AI companies launch new products. Checking a maintained source such as the ai.robots.txt GitHub project quarterly is a reasonable minimum for most UK businesses.
Is Bytespider dangerous and should I block it by default?
Bytespider, operated by ByteDance, has been reported to disregard robots.txt instructions in some cases, making standard blocking rules unreliable against it. It ranked among the most-blocked crawlers in Q1 2026 robots.txt data, appearing in 4.23% of DISALLOW rules, according to Technology Checker via Digital Applied, so WAF-level IP blocking is the more reliable control.
Tracking your AI crawler visibility with Aether AI
Deciding which AI crawlers to allow is only half the problem; the harder question is whether blocking or allowing actually changes your citation share in ChatGPT, Perplexity, Gemini, Claude, Copilot and Google AI Overviews. Aether AI addresses this directly through citation tracking across all six major AI engines, alongside keyword and competitor tracking and Google Search Console integration, so a UK business can see the real effect of a robots.txt change rather than guessing.
Aether AI's platform is the same engine the agency runs for its own clients, including Priority First, Aether Agency and Pulse Operations — the tool is proven on live accounts before it reaches self-serve users.
If you're unsure whether your current crawler settings are helping or hurting your AI visibility, run a free AI-visibility audit at /audit to see where your site currently stands across the major engines.