Mozilla-Tabstack
Mozilla-Tabstack is a bot operated by Mozilla that crawls pages into a dataset that is then reused or sold. Tabstack is a web intelligence API for AI agents. It is documented as honouring robots.txt, but no IP range list or reverse-DNS convention is published, so a request carrying this user agent cannot be proven genuine.
Mozilla-Tabstack at a glance
| User agent | Mozilla-TabstackToken only. The operator has not published a full user-agent string, so match on the substring rather than an exact header. |
|---|---|
| Operator | Mozilla |
| Purpose | Data scrapers and resellers |
| Respects robots.txt | Yes, documented |
| Verification method | None published, user agent only No IP range file and no reverse-DNS convention have been published for this agent. The user-agent string is the only signal, and any client can send it, so treat a match as a claim rather than proof. |
| Crawl pattern | On demand via API. |
| Robots.txt tokens | Mozilla-Tabstack |
| Operator documentation | None published |
Tabstack is a web intelligence API for AI agents. It extracts structured data from web pages and makes it available to AI agents.
Allow or block Mozilla-Tabstack
Paste one of these into the robots.txt file at the root of your domain. Rules are per token, so a block on one crawler leaves every other bot untouched.
User-agent: Mozilla-Tabstack
Disallow: /Blocks every path for this crawler only.
User-agent: Mozilla-Tabstack
Allow: /Explicit allow. Useful when a wildcard rule above it would otherwise catch this crawler.
User-agent: Mozilla-Tabstack
Allow: /
Disallow: /account/
Disallow: /checkout/
Disallow: /searchEdit the Disallow paths to match your own account, checkout and search URLs.
What blocking actually costs you
Blocking Mozilla-Tabstack stops your content being packaged and resold to third parties. Whoever bought earlier crawls still has them.
Block every crawler in the directory at onceIs that really Mozilla-Tabstack?
A user-agent header is a string the client chooses. Scrapers copy Mozilla-Tabstack precisely because site owners allow it. No IP range file and no reverse-DNS convention have been published for this agent. The user-agent string is the only signal, and any client can send it, so treat a match as a claim rather than proof.
Paste the IP address from your access log below. You will see whether it belongs to a hosting provider, a residential proxy pool or a Tor exit, plus the ASN that announces it, which is what tells you whether the claim holds up.
Related crawlers
Identify crawlers automatically instead of by hand
Agentscan checks a request against published crawler ranges, reverse DNS and hosting data, then returns a verdict your edge can act on. One call per request, no range files to keep up to date.