Last updated: 28 September 2026
AI Crawlability Checklist: A UK Business Guide to Getting Found by AI Search
An AI crawlability checklist is a set of technical checks — robots.txt rules, rendering, structured data and monitoring — that confirms AI bots like GPTBot and ClaudeBot can access and cite a website. GPTBot's share of verified bot traffic rose from 4.7% to 11.7% between July 2026 and July 2026, making this a growing priority for every UK business.
Key Takeaways
- Aether AI found GPTBot's share of verified bot traffic rose from 4.7% in July 2026 to 11.7% in July 2026, according to Cloudflare via PPC Land.
- Only 37% of the top 10,000 domains on Cloudflare have a robots.txt file, and among those, GPTBot is disallowed in just 7.8% of cases, per Cloudflare.
- Aether AI found 36.4% of crawler requests were rejected with a 4xx status in July 2026, with deliberate 403 blocks alone accounting for 21.4% of all crawler requests worldwide, according to the SEOmator AI Crawler Report.
- 53% of pages cited by AI systems have schema markup implemented — roughly three times the rate of uncited pages — per an Ahrefs study reported via Youmind.
- GPTBot is the most-blocked AI crawler in robots.txt directives, appearing in 5.52% of disallow rules in Q1 2026, according to Technology Checker.
What is AI crawlability?
AI crawlability is the technical property of a website that determines whether AI crawlers — automated bots run by companies like OpenAI, Anthropic and Perplexity — can access, parse and reuse its content in AI-generated answers. Traditional SEO crawlability asks whether Googlebot can index a page for search rankings; AI crawlability asks whether a separate fleet of bots, with different user-agents and different rules, can read that same page well enough to quote it.
The distinction matters because these are not the same bots, and blocking one does not block the other. A site can rank perfectly on Google while being entirely invisible to ChatGPT, Perplexity or Google AI Overviews if its robots.txt file or JavaScript rendering shuts AI crawlers out. Vercel's network data shows GPTBot generated 569 million requests in a single month, with Anthropic's Claude adding a further 370 million — together representing roughly 20% of Googlebot's 4.5 billion requests over the same period, according to Vercel. AI crawlers are now a material share of total bot traffic, not a fringe concern.
Which AI crawlers and bots should a UK website allow or block?
AI crawlers fall into two functional categories: bots that fetch content live to answer a specific user query, and bots that harvest content in bulk to train future models. The decision to allow or block each one should be made deliberately, not left to a default robots.txt template.
The main crawlers UK site owners will encounter include:
- GPTBot — OpenAI's training crawler, documented at OpenAI's developer documentation
- OAI-SearchBot and ChatGPT-User — OpenAI's live-retrieval bots that fetch pages in response to a ChatGPT query
- ClaudeBot and anthropic-ai — Anthropic's crawlers for Claude
- PerplexityBot — Perplexity's retrieval crawler
- Google-Extended — the directive controlling whether Google can use content to train Gemini and power AI Overviews, separate from standard Googlebot indexing
- Bytespider — ByteDance's crawler, often blocked more aggressively than others
- CCBot — Common Crawl's bot, which feeds many third-party AI training datasets
GPTBot appears in 5.52% of all robots.txt disallow rules tracked in Q1 2026, ahead of CCBot at 5.08%, ClaudeBot at 4.88%, Google-Extended at 4.44% and Bytespider at 4.23%, according to Technology Checker. News publishers are especially cautious: 79% block at least one AI training bot, yet only 46% block Google-Extended, per the same Digital Applied analysis — suggesting many publishers treat Google's AI features differently from OpenAI's or Anthropic's, likely because Google-Extended blocking can also affect AI Overviews visibility.
For most commercial UK websites seeking AI-driven visibility, the sensible default is to allow retrieval bots (OAI-SearchBot, ChatGPT-User, PerplexityBot) so content can be cited in live answers, while making a considered, documented choice about training bots (GPTBot, ClaudeBot, Bytespider) based on content licensing appetite.
How do I check my robots.txt file for AI crawler rules?
Checking a robots.txt file means visiting yourdomain.co.uk/robots.txt in a browser and reading the User-agent and Disallow lines for each named AI bot. Only 37% of the top 10,000 Cloudflare domains have a robots.txt file at all, according to Cloudflare.
Look for entries like User-agent: GPTBot followed by Disallow: / (a full block) or Disallow: with nothing after it (allowed). If GPTBot, ClaudeBot or PerplexityBot don't appear at all, they are allowed by default under the standard robots.txt convention.
A second, easily missed problem is CDN or firewall-level blocking. Cloudflare's Bot Fight Mode and similar WAF settings can silently reject AI crawlers even when robots.txt permits them — this is why 36.4% of crawler requests were rejected with a 4xx status in July 2026, and deliberate 403 responses alone made up 21.4% of all crawler requests globally, according to the SEOmator AI Crawler Report. A robots.txt audit is incomplete without also checking server logs or CDN dashboards for these silent rejections.
What technical steps make content readable by AI crawlers?
Technical AI-readability means ensuring a page's core content and facts exist in the raw HTML response, not solely rendered client-side via JavaScript after the page loads. Many AI crawlers execute limited or no JavaScript, so content injected by React, Vue or similar frameworks after initial load can be functionally invisible to them, even though a human visitor sees it fine in a browser.
Priority technical fixes include:
| Issue | Why it matters for AI crawlers | Fix |
|---|---|---|
| JavaScript-rendered content | Many AI bots don't execute JS fully | Server-side render or pre-render key content |
| Missing or thin robots.txt | Bots default to "allowed" but firewalls may still block | Add explicit User-agent rules for each bot |
| Content behind logins/paywalls | Bots cannot authenticate | Offer a crawlable summary or excerpt |
| Slow server response | Bots time out and abandon the crawl | Improve TTFB and hosting performance |
| No canonical tags | Bots may cite duplicate or wrong URLs | Set clear self-referencing canonicals |
| No llms.txt file | Bots can't find a curated content map | Publish an llms.txt per the llms.txt specification |
The llms.txt standard, proposed by Jeremy Howard and covered by Search Engine Land, gives AI systems a plain-text map of a site's most important pages — separate from, and complementary to, robots.txt.
"An AI crawlability checklist should start with access, not content: check robots rules allow the major assistant crawlers, confirm pages render without heavy scripting, and make sure canonical tags point somewhere sensible. Then look at structure — clear headings, one answer per page, no vital fact buried after a scroll. If a crawler can't reach the page or can't parse it, the best writing in the world goes unseen." — Lauren Dawkins, Head of Content, Aether AI
How does schema markup help AI systems cite content correctly?
Schema markup is structured data code — typically written in JSON-LD format — that labels a page's content in machine-readable terms, such as marking a paragraph as a FAQPage answer or a product's price as a Product schema field. It gives AI systems an explicit, unambiguous signal about what a piece of content is, rather than requiring them to infer it from surrounding text.
The correlation between schema and AI citation is strong: 53% of pages cited by AI systems have schema markup implemented, roughly three times the implementation rate seen on pages that are never cited, according to an Ahrefs study reported via Youmind. However, the same research is more cautious about causation — a controlled Ahrefs test of 1,885 pages found no significant lift in citations purely from adding schema. The honest reading is that schema markup correlates with the kind of well-structured, well-maintained sites that AI systems already favour, rather than acting as a standalone citation trigger on its own.
For UK businesses, the practical takeaway is to treat schema as good hygiene alongside — not instead of — clear writing, accurate facts and clean HTML. Organization, Article, FAQPage and Product schema types remain the most relevant for most commercial sites.
What common mistakes make websites invisible to AI answer engines?
The most common mistake is assuming AI crawlability equals traditional SEO — a site can rank on page one of Google while being entirely absent from ChatGPT or Perplexity answers, because the crawlers and ranking logic differ. Businesses often audit only Google Search Console and never check server logs for GPTBot, ClaudeBot or PerplexityBot activity at all.
Other frequent failures include:
- Blanket-blocking all bots via a security plugin or WAF rule, without realising it also blocks legitimate AI retrieval crawlers alongside malicious scrapers
- Confusing robots.txt permission with actual access — a page can be "allowed" in robots.txt yet still return a 403 at the server or CDN level
- Publishing content that only exists as an image, PDF or embedded video with no accompanying text that a crawler can parse
- Never verifying bot authenticity, allowing spoofed user-agents claiming to be GPTBot to bypass security under a false identity
- Treating AI crawlability as a one-off project rather than an ongoing discipline, as bot behaviour and standards change quickly — GPTBot's traffic share alone more than doubled in twelve months, per Cloudflare
A site that ticks every box once in January and never checks again is likely to drift out of compliance with new crawler rules within months.
Who is responsible for AI crawlability within a business?
AI crawlability responsibility typically spans three roles: developers who control robots.txt, server configuration and rendering; SEO or marketing leads who decide crawler policy and monitor citation performance; and content teams who structure pages for extraction. No single department owns it end-to-end, which is precisely why it gets neglected.
A practical split for a UK business: developers implement and maintain robots.txt, llms.txt and server-side rendering fixes; the SEO or marketing lead sets policy on which bots to allow, tracks citation appearances across AI engines, and reports results to stakeholders; content writers ensure each page opens with a clear, extractable answer to its core question. Agencies managing multiple client sites should document these decisions per domain, since crawler policy is rarely identical across a portfolio — a media client protecting licensed content will make different choices to an e-commerce brand chasing visibility.
Aether AI's own citation tracking, covering ChatGPT, Perplexity, Google AI Overviews, Claude, Gemini and Copilot, exists precisely because this ownership gap causes AI visibility work to fall through the cracks between departments.
Your AI crawlability checklist
- Check robots.txt at yourdomain.co.uk/robots.txt for explicit rules covering GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and Google-Extended.
- Test server and CDN-level blocking by reviewing logs for 403 or 429 responses to known AI bot user-agents.
- Confirm core content renders without JavaScript by disabling JS in a browser and checking the page still shows key facts.
- Add or update an llms.txt file listing your most important pages per the llms.txt specification.
- Implement schema markup for Organization, Article and FAQPage types across key pages.
- Set canonical tags on every page to prevent duplicate-URL confusion.
- Verify bot authenticity via official IP ranges before blocking suspicious traffic labelled as an AI crawler.
- Review and repeat this checklist quarterly, given how fast crawler policy and bot volumes are changing.
FAQ
What does AI crawlability mean?
AI crawlability is whether AI bots — such as GPTBot, ClaudeBot or PerplexityBot — can access, read and reuse a website's content in AI-generated answers. It differs from traditional SEO crawlability because it depends on separate crawlers, separate robots.txt rules, and separate rendering requirements to those used by Googlebot for search rankings.
Does blocking GPTBot affect my Google search rankings?
No, blocking GPTBot does not affect standard Google search rankings, because GPTBot and Googlebot are entirely separate crawlers controlled by separate robots.txt directives. However, blocking Google-Extended specifically can affect whether your content appears in Google AI Overviews, since that directive governs Google's AI training and generative features rather than core search indexing.
What is the difference between robots.txt and llms.txt?
Robots.txt is a decades-old standard that tells crawlers which parts of a site they may or may not access, using Allow and Disallow rules per user-agent. The llms.txt specification is a newer, complementary file that gives AI systems a curated, plain-text map of a site's most important pages, helping them find and prioritise key content rather than simply gating access.
How do I check if my site is blocking AI crawlers?
Visit yourdomain.co.uk/robots.txt directly in a browser and look for User-agent entries naming GPTBot, ClaudeBot, PerplexityBot or Google-Extended, then check whether Disallow rules block them. It's also worth checking server or CDN logs for 403 responses, since firewall-level blocking can silently override robots.txt permissions.
Does schema markup actually improve AI citations?
Schema markup correlates strongly with AI citation, with 53% of AI-cited pages carrying schema versus a much lower rate among uncited pages, according to an Ahrefs study reported via Youmind. A controlled test in the same research found no significant standalone causal lift from adding schema alone, suggesting it works best alongside strong content structure rather than as a fix on its own.
How can I tell if AI assistants are citing my website?
The most direct method is manually querying ChatGPT, Perplexity, Google AI Overviews and other engines with questions relevant to your content and checking whether your site is named or linked. For ongoing measurement rather than manual spot-checks, dedicated citation tracking tools monitor mentions across multiple AI engines automatically and report changes over time.
How much does an AI crawlability audit typically cost?
Costs vary widely depending on site size and whether the work is done in-house, via an agency, or through self-service software — a simple robots.txt and llms.txt review can take a developer a few hours, while a full technical audit across a large or JavaScript-heavy site involves considerably more work. Aether AI offers a free AI-visibility audit at /audit as a starting point before committing to paid tools or agency work.
Are there GDPR or data protection issues with allowing AI bots to crawl my site?
AI crawlers accessing publicly published content generally raise fewer GDPR concerns than crawlers accessing personal data, but businesses should still review what personal information appears on crawlable pages, since the UK GDPR and the Information Commissioner's Office (ICO) guidance on web scraping and AI training data continue to evolve. Sites containing customer data, case studies with identifiable individuals, or user-generated content should apply the same data minimisation principles to AI crawlers as to any other automated access.
Getting AI-crawlable with Aether AI
Every issue covered in this checklist — blocked bots, JavaScript rendering gaps, missing schema, invisible pages — has one thing in common: they're invisible until someone measures them. Aether AI was built to close that gap, running the same technical GEO diagnostics and citation tracking that Aether Agency Ltd uses for its own clients, including Priority First, Aether Agency and Pulse Operations.
Aether AI tracks citation appearances across six AI engines — ChatGPT, Perplexity, Google AI Overviews, Claude, Gemini and Copilot — alongside keyword and competitor tracking and Google Search Console integration, giving businesses a single view of where AI crawlability is working and where it's silently failing.
Run a free AI-visibility audit at /audit to see exactly which AI crawlers can reach your site today, and check Aether AI's public pricing to find the plan that fits your team.