# Visitor classification

Ruleset version: `rules-2026.09.30`, registry updated: 2026-09-30.

## Five classes (mutually exclusive)

| actor_class | Meaning |
|---|---|
| human_likely | Likely human, supported by multiple browser and behavioral signals; not proof of a human |
| ai_verified | Known AI client + independent identity verification (official IP list or trusted platform signal) |
| ai_claimed | Only an AI client UA claim, unverified |
| other_automation | Search crawlers, monitoring, scanners, command-line tools, automation frameworks, etc. |
| unknown | Insufficient evidence; kept as is rather than forced into a class |

## Principles

- A spoofed UA such as GPTBot cannot become ai_verified; reports from the browser SDK cannot carry any trusted network evidence.
- verifiedBot=true does not mean AI; ordinary search crawlers are other_automation.
- No JS does not mean a bot; an automated browser that runs JS is not automatically counted as human.
- Missing Cloudflare fields are recorded as null / unsupported, not treated as a score of 0.
- Google-Extended is only a robots.txt control token, not a separate UA.
- AI referrals (channel=ai_referral) are matched exactly on the referrer domain and are independent of visitor classification; AI crawling and human visits referred by AI are counted separately.

## Client registry

| Provider | Client | Category | Purpose | Verification | Official source |
|---|---|---|---|---|---|
| OpenAI | GPTBot | ai | training | ip_list | [Link](https://developers.openai.com/api/docs/bots) |
| OpenAI | OAI-SearchBot | ai | search | ip_list | [Link](https://developers.openai.com/api/docs/bots) |
| OpenAI | ChatGPT-User | ai | user_triggered | ip_list | [Link](https://developers.openai.com/api/docs/bots) |
| OpenAI | OAI-AdsBot | ai | monitoring | ip_list | [Link](https://developers.openai.com/api/docs/bots) |
| Anthropic | ClaudeBot | ai | training | ip_list | [Link](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) |
| Anthropic | Claude-SearchBot | ai | search | ip_list | [Link](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) |
| Anthropic | Claude-User | ai | user_triggered | ip_list | [Link](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) |
| Perplexity | PerplexityBot | ai | search | none | [Link](https://docs.perplexity.ai/guides/bots) |
| Perplexity | Perplexity-User | ai | user_triggered | none | [Link](https://docs.perplexity.ai/guides/bots) |
| Google | Googlebot | search | search_indexing | ip_list | [Link](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers) |
| Google | Google-Extended | ai | training | none | [Link](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers) |
| Microsoft | bingbot | search | search_indexing | ip_list | [Link](https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0) |
| Baidu | Baiduspider | search | search_indexing | none | [Link](https://help.baidu.com/question?prod_id=99&class=476&id=2996) |
| Yandex | YandexBot | search | search_indexing | none | [Link](https://yandex.com/support/webmaster/robot-workings/check-yandex-robots.html) |
| Sogou | Sogou web spider | search | search_indexing | none | [Link](https://www.sogou.com/docs/help/webmasters.htm) |
| UptimeRobot | UptimeRobot | monitoring | monitoring | none | [Link](https://uptimerobot.com/) |
| SolarWinds | Pingdom | monitoring | monitoring | none | [Link](https://www.pingdom.com/) |
| Ahrefs | AhrefsBot | seo | scraping | none | [Link](https://ahrefs.com/robot) |
| Semrush | SemrushBot | seo | scraping | none | [Link](https://www.semrush.com/bot/) |
| Meta | facebookexternalhit | preview | unknown | none | [Link](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/) |
| Slack | Slackbot-LinkExpanding | preview | unknown | none | [Link](https://api.slack.com/robots) |
| X | Twitterbot | preview | unknown | none | [Link](https://developer.x.com/en/docs/x-for-websites/cards/guides/getting-started) |
| zmap | zgrab | scanner | scanning | none | [Link](https://github.com/zmap/zgrab2) |
| ProjectDiscovery | Nuclei | scanner | scanning | none | [Link](https://github.com/projectdiscovery/nuclei) |
| sqlmap | sqlmap | scanner | scanning | none | [Link](https://sqlmap.org/) |
| curl | curl | cli | tooling | none | [Link](https://curl.se/) |
| GNU | Wget | cli | tooling | none | [Link](https://www.gnu.org/software/wget/) |
| Python | python-requests | cli | tooling | none | [Link](https://requests.readthedocs.io/) |
| Python | python-httpx | cli | tooling | none | [Link](https://www.python-httpx.org/) |
| Go | Go-http-client | cli | tooling | none | [Link](https://pkg.go.dev/net/http) |
| Node.js | node-fetch/undici | cli | tooling | none | [Link](https://nodejs.org/) |
| Google | HeadlessChrome | automation | unknown | none | [Link](https://developer.chrome.com/docs/chromium/headless) |
| PhantomJS | PhantomJS | automation | unknown | none | [Link](https://phantomjs.org/) |
