An AI crawler is a bot that visits websites on behalf of an artificial intelligence company. It can do so for three different reasons: to collect text for training a model, to index pages for an AI search engine, or to open a page in real time because a user asked for it in a chat. Which ones you let in depends on which of those three you want to allow.
We counted the bots that visited visilay.com between 31 August and 27 September 2026, across roughly 199,000 lines of logs. OpenAI's bots (GPTBot, OAI-SearchBot and ChatGPT-User together) made 4,773 requests. Googlebot made 1,355.
A month of requests on our own site
| User agent | Company | Requests (31 Aug - 27 Sep 2026) |
|---|---|---|
| Claude-User | Anthropic | 9,340 |
| GPTBot | OpenAI | 2,689 |
| bingbot | Microsoft | 2,281 |
| Applebot | Apple | 1,567 |
| ClaudeBot | Anthropic | 1,540 |
| Googlebot | 1,355 | |
| Meta-ExternalAgent | Meta | 1,353 |
| OAI-SearchBot | OpenAI | 1,250 |
| ChatGPT-User | OpenAI | 834 |
| GoogleOther | 617 | |
| Amazonbot | Amazon | 107 |
| PerplexityBot | Perplexity | 100 |
| Perplexity-User | Perplexity | 58 |
| Bytespider | ByteDance | 57 |
| CCBot | Common Crawl | 32 |
| Claude-SearchBot | Anthropic | 30 |
| MistralAI-User | Mistral | 29 |
| DuckAssistBot | DuckDuckGo | 24 |
Three things to bear in mind when reading the table:
- Claude-User is inflated by us. We use Claude every day to work on the site, so a good share of those requests are our own. Don't read them as outside interest.
- ChatGPT-User visited 70 different URLs, but 402 of its 834 requests were for the home page. That's typical of an assistant checking who we are before answering a question.
- Requests don't mean human visits. Over the same period, visits with a chatgpt.com referrer, meaning people who clicked a link in an answer, came to 18. Our site's audience is mostly Italian, so the mix of bots and referrals on a UK site may look different.
The underlying pattern matches what Cloudflare measures. In its July 2026 report, training crawlers make up 52% of AI bot traffic, up from 22% in spring 2025. A year earlier Cloudflare had calculated that Anthropic made around 70,900 requests for every visit it referred back to a site. What the bots take and what they give back in traffic are badly out of balance.
The three types of AI crawler
Training crawlers
They download pages to train models. Blocking them doesn't remove your site from today's answers, but it reduces how much future models will know about you.
- GPTBot (OpenAI): blocking it is the same as opting out of training (documentation).
- ClaudeBot (Anthropic).
- Google-Extended: not a separate crawler but a robots.txt token. It controls whether content is used to train and ground Gemini, and according to Google it "does not impact a site's inclusion in Google Search".
- CCBot (Common Crawl), Meta-ExternalAgent, Bytespider, Amazonbot.
AI search crawlers
They index pages so they can show them in answers that use web search. Blocking them removes your site from those answers.
- OAI-SearchBot (OpenAI): if you block it, your site won't appear in ChatGPT search summaries. At most, OpenAI can show a link and a title if the page comes from a third-party search provider.
- PerplexityBot: indexes for search, not for training, and respects robots.txt (documentation).
- Claude-SearchBot (Anthropic).
- bingbot: not an AI crawler, but it feeds Copilot and, according to independent tests, some of ChatGPT's results.
User-initiated agents
They open a page because someone asked for it at that moment, for example by pasting a link into a chat or asking for a comparison of suppliers.
- ChatGPT-User: according to OpenAI, "robots.txt rules may not apply".
- Perplexity-User: "generally ignores robots.txt".
- Claude-User, MistralAI-User, DuckAssistBot.
- Google-Agent: added to Google's documentation on 20 March 2026 for Google agents that browse and act at a user's request.
What happens if you block them
| If you block | Effect |
|---|---|
| GPTBot, ClaudeBot, CCBot | Future models are trained without your text. No immediate effect on answers that use search |
| Google-Extended | No effect on Google Search. Limits the use of your content for Gemini |
| OAI-SearchBot | Your site drops out of ChatGPT search answers |
| PerplexityBot | Your site drops out of Perplexity answers |
| ChatGPT-User, Perplexity-User | Often nothing via robots.txt, because they don't always respect it. Via a firewall you also block real customers using an assistant |
| Googlebot | Your site drops out of Google, AI Overviews and AI Mode included |
The last row needs saying plainly: there's no way to stay in Google Search and drop out of AI Overviews by blocking a bot. Since June 2026 Search Console has had a setting to exclude a site from AI features, at domain level. We cover it in how AI Overviews affect SEO.
What we recommend for a business website
For a B2B company or an online shop that wants to be found, the default choice is:
- let in the AI search crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot) and bingbot;
- let in user-initiated agents, because they're potential customers asking a question about you;
- decide case by case on training crawlers. For a company that wants models to know about it, letting them in makes sense. For a publisher that lives on original content, less so.
A sample robots.txt that blocks only OpenAI's training crawler and leaves everything else open:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: *
Allow: /
The full syntax is in our guide to robots.txt.
Firewalls and anti-bot protection
Plenty of sites block AI crawlers without knowing it, because their firewall or CDN treats them as suspicious traffic. It's the most common case we find in audits: robots.txt open, but requests turned away with a 403 error.
There are two ways to tell genuine agents from those pretending to be them:
- published IP lists. OpenAI and Perplexity publish their bots' addresses in JSON format;
- request signatures (Web Bot Auth). Cloudflare introduced "signed agents" in August 2025. ChatGPT's agent signs its requests with the header
Signature-Agent: "https://chatgpt.com".
If you use Cloudflare, check its rules on AI bots. Since July 2025 Cloudflare has blocked AI crawlers by default on new domains, so if your site was set up after that date they may be shut out without you knowing.
How to read your own logs
- Download your access logs from your hosting control panel. In cPanel they're under "Raw Access".
- Search for the user agents in the table above. Even a simple
grep -c "GPTBot" access.loggives you a first number. - Check the response codes. If a bot you want to let in is getting lots of 403s or 429s, something is blocking it.
- Look at which pages they visit. If OAI-SearchBot only visits the home page, your internal pages are hard to reach.
- Compare with human traffic from AI in GA4's AI Assistant channel, as explained in AI visibility.
User agents can be faked. For a serious analysis, check the IPs against the official lists.
Further reading: SEO for AI agents: how to make your site usable by browsers and agents.
FAQs
It's a bot that visits websites on behalf of an artificial intelligence company. It can collect text for training, like GPTBot, index pages for answers that use search, like OAI-SearchBot, or open pages in real time at a user's request, like ChatGPT-User.
No. GPTBot is used for training. For answers that use web search, ChatGPT uses OAI-SearchBot: block that one and your site drops out of ChatGPT search summaries.
Not with robots.txt, because AI Overviews and AI Mode use Googlebot. Since June 2026 Search Console has offered a setting to exclude a site from AI features, which applies to the whole domain.
Because they download a lot of pages and send few visitors back. In 2025 Cloudflare calculated that Anthropic made around 70,900 requests for every visit it referred to a site. On our own site, over four weeks, OpenAI's bots made 4,773 requests and ChatGPT sent 18 visits.