AntileakBot

If you found this page in your server logs, this explains exactly what was requested and why — and how to stop it.

AntileakBot is the crawler behind antileak.io, a service that checks whether AI assistants such as ChatGPT, Claude and Perplexity can read a business's website. It identifies itself as:

Mozilla/5.0 (compatible; AntileakBot/1.0; +https://antileak.io/bot)

When it visits

Only when a person asks it to. A scan is triggered by someone entering a domain on our homepage — there is no background crawl, no discovery queue, and no re-crawling on a schedule. One request from a visitor produces one scan.

The same address can be scanned at most once every three minutes, so AntileakBot cannot become a source of sustained load even if someone tries.

What it requests

RequestWhy
GET /Read the homepage's raw HTML — the same bytes an AI crawler gets, since none of them run JavaScript.
GET /robots.txtCheck whether AI crawlers are disallowed.
HEAD /llms.txtCheck whether the site publishes one.
HEAD on up to 8 internal linksFind broken links reachable from the homepage.
HEAD /sitemap.xmlCheck whether the site publishes one — and only when your robots.txt does not already say. A Sitemap: line there answers the question, and then no request is made at all.

That is the whole of it: twelve requests at the very most, all read-only, plus one more if http:// redirects to https://. Thirteen is a hard limit enforced in code, not a description of what usually happens — a scan carries a budget of thirteen requests and stops when it runs out, so redirects and retries come out of the same thirteen rather than adding to them. Requests are made one at a time, in that order, so the most we ever have open against your server is one.

AntileakBot does not execute JavaScript, does not log in, does not submit forms, does not follow links beyond the homepage, and never requests the same page twice in a scan. It also looks up your domain's public DNS TXT records (SPF, DKIM, DMARC) through Cloudflare's resolver, which does not touch your server at all.

If your server pushes back

Any of these ends the scan immediately. Whatever was left in the list above is abandoned, and the person who asked for the scan is told it was stopped — they are not given results we didn't collect.

Your responseWhat AntileakBot does
429 Too Many RequestsStops. It does not retry — retrying a 429 is pushing a server that has just asked to be pushed less.
403 ForbiddenStops. A refusal is a refusal, whether it comes from robots.txt, your WAF or your firewall.
503 Service UnavailableWaits and retries at most twice (0.5s, then 1s), honouring Retry-After when it asks for three seconds or less. A longer Retry-After, or a third 503, stops the scan.
It does not keep a copy of your pages. A scan stores its results — the scores, which checks passed, which were missing — along with your homepage's <title> and meta description, because those are what the report quotes back to you. The page body is measured and discarded. Nothing is used to train a model.

How to block AntileakBot

Add this to your robots.txt. AntileakBot honours it, and a blocked scan simply reports that it was blocked:

User-agent: AntileakBot
Disallow: /

Blocking by user agent at your firewall or CDN also works, and so does a 403 or a 429 from anywhere — see above.

Contact

Questions, or want your domain permanently excluded from scanning? Email services@antileak.io and we will add it to a suppression list that applies before any request is made.