FacebookBot
FacebookBot is a bot operated by Meta/Facebook that collects pages used to train AI models. Officially used for training Meta "speech recognition technology," unknown if used to train Meta AI specifically. It honours robots.txt, and a request claiming to be FacebookBot can be checked against the operator's announced ASN.
FacebookBot at a glance
| User agent | FacebookBotToken only. The operator has not published a full user-agent string, so match on the substring rather than an exact header. |
|---|---|
| Operator | Meta/Facebook |
| Purpose | Training language models |
| Respects robots.txt | Yes, documentedsource |
| Verification method | Operator ASN check Check the source IP against AS32934. Meta publishes no per-bot range file. https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/ |
| Crawl pattern | Up to 1 page per second |
| Robots.txt tokens | FacebookBot |
| Operator documentation | https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/ |
Officially used for training Meta "speech recognition technology," unknown if used to train Meta AI specifically.
Allow or block FacebookBot
Paste one of these into the robots.txt file at the root of your domain. Rules are per token, so a block on one crawler leaves every other bot untouched.
User-agent: FacebookBot
Disallow: /Blocks every path for this crawler only.
User-agent: FacebookBot
Allow: /Explicit allow. Useful when a wildcard rule above it would otherwise catch this crawler.
User-agent: FacebookBot
Allow: /
Disallow: /account/
Disallow: /checkout/
Disallow: /searchEdit the Disallow paths to match your own account, checkout and search URLs.
What blocking actually costs you
Blocking FacebookBot keeps your pages out of the next training run. It does not remove anything already collected, and it has no effect on how the site ranks in ordinary search.
Block every crawler in the directory at onceIs that really FacebookBot?
A user-agent header is a string the client chooses. Scrapers copy FacebookBot precisely because site owners allow it. Check the source IP against AS32934. Meta publishes no per-bot range file.
Paste the IP address from your access log below. You will see whether it belongs to a hosting provider, a residential proxy pool or a Tor exit, plus the ASN that announces it, which is what tells you whether the claim holds up.
Other crawlers run by Meta/Facebook
Related crawlers
Identify crawlers automatically instead of by hand
Agentscan checks a request against published crawler ranges, reverse DNS and hosting data, then returns a verdict your edge can act on. One call per request, no range files to keep up to date.