Training crawlers

FacebookBot

FacebookBot is a bot operated by Meta/Facebook that collects pages used to train AI models. Officially used for training Meta "speech recognition technology," unknown if used to train Meta AI specifically. It honours robots.txt, and a request claiming to be FacebookBot can be checked against the operator's announced ASN.

FacebookBot at a glance

Reference data for the FacebookBot crawler
User agentFacebookBot

Token only. The operator has not published a full user-agent string, so match on the substring rather than an exact header.

OperatorMeta/Facebook
PurposeTraining language models
Respects robots.txtYes, documentedsource
Verification methodOperator ASN check

Check the source IP against AS32934. Meta publishes no per-bot range file.

https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/
Crawl patternUp to 1 page per second
Robots.txt tokensFacebookBot
Operator documentationhttps://developers.facebook.com/docs/sharing/webmasters/web-crawlers/

Officially used for training Meta "speech recognition technology," unknown if used to train Meta AI specifically.

robots.txt

Allow or block FacebookBot

Paste one of these into the robots.txt file at the root of your domain. Rules are per token, so a block on one crawler leaves every other bot untouched.

Block FacebookBot
User-agent: FacebookBot
Disallow: /

Blocks every path for this crawler only.

Allow FacebookBot
User-agent: FacebookBot
Allow: /

Explicit allow. Useful when a wildcard rule above it would otherwise catch this crawler.

Allow, except the pages you never want quoted
User-agent: FacebookBot
Allow: /
Disallow: /account/
Disallow: /checkout/
Disallow: /search

Edit the Disallow paths to match your own account, checkout and search URLs.

What blocking actually costs you

Blocking FacebookBot keeps your pages out of the next training run. It does not remove anything already collected, and it has no effect on how the site ranks in ordinary search.

Block every crawler in the directory at once
Verification

Is that really FacebookBot?

A user-agent header is a string the client chooses. Scrapers copy FacebookBot precisely because site owners allow it. Check the source IP against AS32934. Meta publishes no per-bot range file.

Paste the IP address from your access log below. You will see whether it belongs to a hosting provider, a residential proxy pool or a Tor exit, plus the ASN that announces it, which is what tells you whether the claim holds up.

Other crawlers run by Meta/Facebook

Related crawlers

Agentscan

Identify crawlers automatically instead of by hand

Agentscan checks a request against published crawler ranges, reverse DNS and hosting data, then returns a verdict your edge can act on. One call per request, no range files to keep up to date.

FAQ

FacebookBot questions

FacebookBot is a bot operated by Meta/Facebook that collects pages used to train AI models. Officially used for training Meta "speech recognition technology," unknown if used to train Meta AI specifically. It honours robots.txt, and a request claiming to be FacebookBot can be checked against the operator's announced ASN.