Totallytics docs
AI crawlers
Totallytics tracks AI crawlers like GPTBot, ClaudeBot and PerplexityBot through a read-only Cloudflare token: 30 days imported on connect, then a sync every 15 minutes, with no code on your site.
Why your tracker can't see them
GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot and the search engine crawlers fetch the raw HTML of a page and never run JavaScript, so the Totallytics script never fires for them and they never show up as visits. Cloudflare sits in front of your site and sees every request, crawlers included. Totallytics reads those requests from Cloudflare's analytics and shows them on the Crawlers tab of your dashboard.
Connect Cloudflare
Your site must be verified in Totallytics and its domain proxied through Cloudflare, since Cloudflare only sees traffic it proxies.
- Open your site's settings, go to Integrations and find Cloudflare.
- Click Create a read-only token. Cloudflare opens its token page with Zone Read and Analytics Read already selected for all your zones. Create the token and copy it.
- Paste it into Paste the token and click Connect.
Connecting starts tracking on every verified site whose domain is in the token's zones. A site matches the zone named after its hostname or a parent domain, so blog.example.com matches the zone example.com. Sites you add later get a Start tracking button on their Crawlers tab.
Only requests to the site's own hostname and its www twin count: example.com also counts www.example.com, but not other subdomains in the same zone.
Totallytics checks the token before saving it and stores it encrypted. It only ever shows the last 4 characters. One token covers every site you own; Replace token swaps it and Disconnect removes it.
What counts as verified
A hit is verified when Cloudflare verified the bot, in its AI Crawler, AI Search, AI Assistant or Search Engine Crawler category, or when the request came from an IP range the bot's company publishes. Totallytics checks the lists of OpenAI, Perplexity, Google, Microsoft, Apple, Common Crawl and DuckDuckGo. The IP check happens during the sync and IP addresses are never stored.
Cloudflare no longer verifies PerplexityBot. On our own sites, about 90% of the PerplexityBot hits Cloudflare didn't verify came from Perplexity's published ranges, so those count as verified. Anthropic publishes no ranges, so Claude's bots only count as verified when Cloudflare verifies them.
A request that claims to be a known bot, like GPTBot or Googlebot, without being verified counts as unverified. Those hits are kept apart from the verified numbers, because anyone can put GPTBot in a user agent: on our sites, unverified hits claiming OAI-SearchBot, GPTBot, ChatGPT-User or Googlebot came from IPs outside those companies' published ranges.
SEO tools, webhooks, link previews, accessibility, advertising and security bots are left out.
Bots we recognize
| Bot | Company | Type |
|---|---|---|
| GPTBot | OpenAI | AI training |
| OAI-SearchBot | OpenAI | AI search |
| ChatGPT-User | OpenAI | AI assistant |
| ClaudeBot | Anthropic | AI training |
| Claude-SearchBot | Anthropic | AI search |
| Claude-User | Anthropic | AI assistant |
| PerplexityBot | Perplexity | AI search |
| Perplexity-User | Perplexity | AI assistant |
| Googlebot | Search engine | |
| GoogleOther | AI training | |
| Bingbot | Microsoft | Search engine |
| Applebot | Apple | AI search |
| meta-externalagent | Meta | AI training |
| meta-externalfetcher | Meta | AI assistant |
| Amazonbot | Amazon | AI training |
| Bytespider | ByteDance | Search engine |
| CCBot | Common Crawl | AI training |
| DuckDuckBot | DuckDuckGo | Search engine |
| DuckAssistBot | DuckDuckGo | AI assistant |
| MistralAI-User | Mistral | AI assistant |
| ExaSearchBot | Exa | AI search |
| Baiduspider | Baidu | Search engine |
Googlebot includes Googlebot-Image, Googlebot-Video and Googlebot-News. Other bots Cloudflare verifies, such as PetalBot or YandexUserproxy, appear under the name in their user agent, with the type of their Cloudflare category. A verified request with a generic user agent shows as Unidentified.
What the numbers mean
- AI crawler hits: verified hits from AI search, AI assistant and AI training bots.
- Search engine hits: verified hits from search engine crawlers.
- Pages crawled: distinct paths with at least one verified hit.
- Unverified hits: hits from requests that claim a known bot but aren't verified.
Each total is compared with the period of the same length right before your range. The chart splits verified hits by type; search engines start hidden, click them in the legend to show them.
The Bots panel lists every bot with its verified hits and when it was last seen, the latest hour with a verified hit. A blocked count shows verified hits your site answered with 401, 403 or 429, so a firewall rule or bot setting that turns crawlers away is easy to spot. An unverified count shows the hits that only claimed to be that bot. Pages lists the top 100 paths with how many different bots read each one, and Responses shows the status codes crawlers got. Click a bot or a page to filter the whole tab.
Hit counts are Cloudflare's sample-adjusted estimates.
Sync and history
| Limit | Value |
|---|---|
| Sync | Every 15 minutes |
| History imported on connect | Last 30 days |
| Cloudflare's own retention | 31 days |
| Totallytics retention | Kept after Cloudflare drops it |
| Granularity | 1 hour |
| Bots and pages per report | Top 100 |
| Cloudflare plan | Any, including Free |
Connecting imports the last 30 days, a week at a time. The Crawlers tab shows the progress and the numbers appear once the import is done. After that, a sync every 15 minutes adds the newest hours and reads the previous hour again to catch late data. Sync now in site settings queues one right away.
Cloudflare keeps this data for 31 days. Totallytics stores it per hour, so your crawler history keeps growing after Cloudflare has dropped it.
If Cloudflare stops accepting the token, the site shows Needs attention with the reason. Replace the token and the sync picks up where it stopped.
Stop tracking and disconnect
Stop tracking in a site's settings stops crawler tracking on that site and deletes its crawler history. Disconnect removes your Cloudflare token, stops tracking on all your sites and deletes their crawler history. Both ask before they delete anything. Tracking a site again imports the last 30 days again.
API
All routes are owner-only. GET /api/sites/{hostname}/crawlers/status and GET /api/sites/{hostname}/crawlers/overview also accept a management key with read, and POST /api/sites/{hostname}/crawlers/sync one with manage. The other routes need the owner's Firebase ID token.
GET,PUTandDELETE /api/cloudflare/connection: read, save ({ "token" }) or remove the Cloudflare token.GET /api/sites/{hostname}/crawlers/status: connection, zone and sync state.POSTandDELETE /api/sites/{hostname}/crawlers/link: start tracking a site, or stop and delete its history.POST /api/sites/{hostname}/crawlers/sync: queue a sync (202).GET /api/sites/{hostname}/crawlers/overview?from&to&tz: totals, previous period, series, bots, pages and status codes.fromandtoare unix seconds,toexclusive. Filter withf=bot:GPTBot,f=vendor:OpenAI,f=purpose:ai_searchorf=path:/pricing, one per key;purposeisai_search,ai_assistant,ai_trainingorsearch.
The full schemas are in the OpenAPI file.
FAQ
Do I need to add code to my site?
No. Totallytics reads Cloudflare's request analytics with a read-only token. Nothing changes on your site or in your Cloudflare settings.
Does it work on the Cloudflare Free plan?
Yes. The analytics Totallytics reads are available on Free zones, with 31 days of history.
Why does Connect reject my token?
The token needs Zone Read to list your zones and Analytics Read to read requests. Connect checks both and says what is wrong: Cloudflare rejected the token, the token can't see any zones, or it can't read analytics.
Why does my site have no Cloudflare zone?
The token can't see a zone for that domain. Create a token that covers it and use Replace token in the site's settings.
What happens to my data when I disconnect?
Totallytics deletes the crawler history of every site that used the token. Your web analytics are not touched.