· 9 min read · Wwwebtech Team
Are AI Crawlers Reaching Your Site? Check Your Logs
GPTBot, ClaudeBot and PerplexityBot either reach your pages or they don't. Here's how to find out from your own server logs, and what blocking really costs.
In this piece
Somewhere in the last two years, a new set of visitors started arriving at your website. They don't fill in forms, they don't buy anything, and they never show up in Google Analytics. They are the crawlers run by the companies behind ChatGPT, Claude, Perplexity and Gemini, and they are reading your pages — or being turned away at the door, often by a setting nobody in your business chose deliberately.
This piece is about three things: which crawlers actually exist, how to check in an afternoon whether they are reaching you, and what you genuinely give up if you block them. On the last point, anyone who gives you a confident answer is guessing. I will tell you where the guessing starts.
Who is actually knocking
A crawler identifies itself with a user agent string — a line of text it sends with every request, saying who it is. Your web server writes that line into a log file. The major AI companies publish their user agent tokens, which means you can look for them by name.
The important thing most people miss: each company runs more than one crawler, and they do different jobs. Blocking all of them with one rule is a blunt instrument.
| Token in robots.txt | Run by | Roughly what it does |
|---|---|---|
| GPTBot | OpenAI | Bulk crawl; OpenAI says this feeds model training |
| OAI-SearchBot | OpenAI | Builds the index behind ChatGPT's search results |
| ChatGPT-User | OpenAI | Fetches a page because a user asked about it, live |
| ClaudeBot | Anthropic | Bulk crawl |
| Claude-User / Claude-SearchBot | Anthropic | Live fetch on a user's request; search indexing |
| PerplexityBot | Perplexity | Indexing for Perplexity's answers |
| Perplexity-User | Perplexity | Live fetch when a user clicks through to a source |
| Google-Extended | Not a crawler. A permission token controlling whether already-crawled pages feed Gemini | |
| CCBot | Common Crawl | Open dataset that many AI models have been trained on |
| Bytespider, Amazonbot, meta-externalagent, Applebot-Extended | ByteDance, Amazon, Meta, Apple | Various mixes of training and product use |
Google-Extended deserves a separate note, because it is the one people get wrong. It is not a bot that visits you. Googlebot still crawls your site for ordinary Search regardless. Google-Extended is a switch in robots.txt that tells Google whether the pages it has already fetched may be used for Gemini and related AI products. Disallowing it does not remove you from Google Search, and allowing it does not get you into Search faster.
What they actually fetch
They fetch HTML. That is the plain text of your page as it leaves your server, before a browser does anything clever with it.
This matters enormously if your site was built as a single-page application, where the server sends an almost-empty shell and JavaScript assembles the content in the visitor's browser. Googlebot handles that reasonably well these days. Several AI crawlers appear not to, or to do it inconsistently. I am deliberately not stating this as fact for any specific bot, because the companies change behaviour without announcement and nobody outside them can see the pipeline. Test it rather than assume it. The test is simple: in Chrome, open your page, press Ctrl+U to view source, and search that raw text for a sentence from your main content. If the sentence isn't there, a non-rendering crawler sees an empty page.
Beyond that, crawlers generally follow links from page to page, respect (or ignore) your robots.txt, and request pages at whatever rate they feel like. A crawler that hammers a slow, shared-hosting site can genuinely cost you — that is a legitimate reason to rate-limit, and a different decision from blocking on principle.
How to check your own logs
Analytics will not help you. Google Analytics runs JavaScript in a browser; crawlers that don't run JavaScript are invisible to it. You need raw server logs — the file your web server writes for every single request.
Where to find them:
- cPanel hosting (most Indian shared hosts): look for "Raw Access" or "Raw Access Logs". You can download a compressed file of recent requests, and usually switch on archiving so old months are kept.
- Cloudflare in front of your site: the Security and Analytics sections have bot traffic breakdowns, and newer accounts get an explicit AI crawler view.
- A VPS or cloud server: typically /var/log/nginx/access.log or /var/log/apache2/access.log.
- WordPress with a security plugin: Wordfence and similar keep a live traffic log that shows user agents.
Each line contains an IP address, a date, the URL requested, an HTTP status code and the user agent. You are looking for two things. First, does the name appear at all? Search the file for "GPTBot", "ClaudeBot", "PerplexityBot". No matches across a month means nothing is reaching you. Second, what status code did they get? 200 means the page was served. 403 means forbidden — something actively refused them. 404 means that page doesn't exist. 429 means too many requests; you or your host are throttling them.
A log full of 403s against GPTBot is the single most common finding, and the owner almost never knew.
Don't trust the name alone
Anyone can send any user agent string. Scrapers routinely pretend to be GPTBot or Googlebot. OpenAI, Anthropic and Perplexity each publish the IP address ranges their crawlers use, so a request claiming to be GPTBot from an IP outside OpenAI's published list is something else wearing a mask. There have also been public disputes — Cloudflare has made allegations about crawlers fetching content from undeclared addresses — which is a good reason to treat user agent strings as a label, not proof.
robots.txt and what it cannot do
Your robots.txt lives at yoursite.in/robots.txt. Type that into a browser right now. It is a plain text file of instructions, and it is a request, not a lock. Well-behaved crawlers obey it. Badly behaved ones read it to find out which directories you'd rather they didn't see.
To let the AI crawlers in explicitly, a group looks like this, one block per crawler:
- User-agent: GPTBot
- Allow: /
To block one:
- User-agent: ClaudeBot
- Disallow: /
Two traps. First, a blanket User-agent: * followed by Disallow: / somewhere in the file blocks everything that doesn't have its own named block — and plenty of sites still carry that line from a staging environment that was never cleaned up when they went live. Second, robots.txt is not the only place blocking happens. A Cloudflare or similar CDN sitting in front of your site can block AI crawlers at the network level, before the request ever reaches your server and before robots.txt is consulted. Cloudflare has introduced one-click AI bot controls and has moved towards blocking AI crawlers by default for some new sites. If you are on Cloudflare, check the dashboard, not just the text file. This is exactly the kind of thing worth having someone look at under ongoing technical support rather than discovering a year later.
What blocking actually costs
Here is where I stop being certain, and so should anyone selling you this.
Separate the question into two. Training is whether your words end up inside a model's weights. You get no attribution, no link, no traffic. The argument for blocking training crawlers is straightforward if you are a publisher whose content is the product. For a dental clinic in Preet Vihar or a components supplier in Jhilmil, there is very little to protect and the upside is that models become faintly more aware your category exists. It is close to a non-decision either way.
Retrieval is different and this is the one that matters. When someone asks ChatGPT or Perplexity for "CCTV installer in east Delhi", these tools fetch live pages and cite them with a clickable link. If you have blocked the retrieval crawlers — OAI-SearchBot, Perplexity-User, Claude-SearchBot — you cannot be cited, full stop. You have removed yourself from a channel that sends real visitors.
How many visitors? Check your own numbers rather than believing anyone's estimate. In Google Analytics 4, look at Traffic Acquisition and filter the session source for chatgpt.com, perplexity.ai and gemini.google.com. For most Indian small businesses today that number is small next to Google. Whether it stays small is genuinely unknown. The sensible posture for most businesses is: allow the retrieval crawlers, make your own call on the training ones, and treat AI visibility as something you measure rather than something you buy.
What I would not pay for
- An llms.txt implementation as a paid deliverable. It's a proposed file format for summarising your site for language models. It's a reasonable idea, it costs nothing to add, and no major AI company has publicly committed to reading it. Paying a meaningful sum for one is paying for a hope.
- "AI ranking" packages with guaranteed placement. There is no ranking to buy. Answers vary between users and between identical queries. Anyone promising position is describing something that does not exist.
- Bot-blocking sold as content protection to a small business. If your site is service pages and a blog, the material being protected is marketing material you want read. The blocking is real; the benefit usually isn't.
- A rewrite of your whole site "for AI". Clean HTML, fast pages, clear headings and factual, well-structured content is the same work that ordinary search has rewarded for a decade. If a proposal for AI readiness looks nothing like good web development, be suspicious.
Do this this week
- Open yoursite.in/robots.txt in a browser and read it. Look for any Disallow: / and work out what it applies to.
- Download a month of raw access logs from your hosting panel and search for GPTBot, ClaudeBot and PerplexityBot. Note whether they appear, and what status code they got.
- If you use Cloudflare or any CDN, check its bot settings — blocking there overrides everything in your robots.txt.
- View the page source of your most important service page and confirm your actual sales copy is in the HTML, not assembled by JavaScript.
- In GA4, filter sessions by source for chatgpt.com and perplexity.ai, so you have a baseline to compare against in six months.
That's an afternoon's work and it replaces speculation with facts about your own site. If the logs throw up something you can't interpret, or you find 403s you can't trace, send us the log extract and your robots.txt and we'll tell you what's blocking what.
Questions we get asked
Will blocking GPTBot hurt my Google rankings?
No. GPTBot is OpenAI's crawler and has nothing to do with Googlebot or Google Search. Blocking it does not affect how Google crawls, indexes or ranks your pages. The separate token Google-Extended also does not affect Search — it only governs whether your content feeds Google's AI products.
Where do I find raw server logs if I'm on shared hosting?
In cPanel, look under Metrics for "Raw Access" or "Raw Access Logs". You can download a compressed file of recent requests and usually enable archiving so previous months are retained. If your host doesn't expose logs at all, a security plugin like Wordfence keeps a live traffic log showing user agents, which is enough for this check.
Can I block AI crawlers from training but still appear in ChatGPT answers?
In principle yes, because the companies run separate crawlers for separate jobs. You would disallow GPTBot and ClaudeBot while allowing OAI-SearchBot, Claude-SearchBot and the user-triggered fetchers. Whether that split holds over time is up to the companies, so it is worth re-checking your robots.txt against their published documentation every few months.
Someone is crawling my site claiming to be GPTBot. Is it really OpenAI?
Not necessarily. Any script can send any user agent string, and scrapers commonly impersonate well-known crawlers. OpenAI, Anthropic and Perplexity each publish the IP ranges their crawlers use, so compare the IP in your log against the published list. A mismatch means it is something else wearing the name.
Do I need an llms.txt file?
It costs nothing to add, so there is no harm in one. But no major AI company has publicly committed to reading it, so treat it as an experiment rather than a requirement, and do not pay a significant fee for one. Clean HTML and clear, factual pages do far more.
If this is your problem
What we’d actually do about it.
Service
Technical SEO & Core Web Vitals
Service