参考

爬虫 API

Agentscan 可验证的爬虫,以及每个爬虫的核验方式。无需密钥。

列出可验证的爬虫

GET/v1/crawlers
无需密钥

响应字段

  • crawlers[].slugstring

    验证时作为 bot 传入的标识符。

  • crawlers[].tokensstring[]

    爬虫发送的 user-agent 标记。

  • crawlers[].methodstring

    爬虫的验证方式,如 ip-list+reverse-dns。

  • crawlers[].verifiableboolean

    运营方公布了可供核对的信息时为 true。

  • crawlers[].reverseDnsstring[]

    真实请求解析到的主机名后缀。

  • crawlers[].prefixFilesstring[]

    检查时使用的已公布 IP 段文件。

  • countinteger

    列出的爬虫数量。

curl https://ipscanner.io/v1/crawlers
响应
{
  "count": 36,
  "crawlers": [
    {
      "slug": "googlebot",
      "name": "Googlebot",
      "tokens": [
        "googlebot"
      ],
      "method": "ip-list+reverse-dns",
      "verifiable": true,
      "reverseDns": [
        "googlebot.com",
        "google.com"
      ],
      "prefixFiles": [
        "https://developers.google.com/static/search/apis/ipranges/googlebot.json"
      ],
      "docs": "https://developers.google.com/search/docs/crawling-indexing/verifying-googlebot"
    },
    {
      "slug": "gptbot",
      "name": "GPTBot",
      "tokens": [
        "gptbot"
      ],
      "method": "ip-list",
      "verifiable": true,
      "prefixFiles": [
        "https://openai.com/gptbot.json"
      ],
      "docs": "https://platform.openai.com/docs/bots"
    }
  ],
  "note": "Crawlers this service can verify by evidence the operator publishes."
}