FigureCheckBot
If this address appeared in your logs, a FigureCheck™ client has asked us to check the risk disclosure on a page of yours, usually because the page promotes their brand. This is what the scanner does, what it does not do, and how to stop it.
User-Agent FigureCheckBot/0.1 (+https://mvrk.systems/figurecheck/bot)
One page at a time,
and only pages it was given.
It does not crawl. Every request is for an address a client named, or a partner page they declared, and it reads that page for one thing: the retail-loss percentage in the disclosure.
- robots.txt first
- Fetched before the page and honoured, including
Crawl-delay. An explicitDisallowis always obeyed; the page is then reported to the client as unreachable rather than read some other way. - One request per host
- At most one request every two seconds to any one host, or slower if your
Crawl-delayasks for it. The limit is per host, so no single site absorbs the scanner's volume. - Backs off when asked
- A
429or a5xxpushes the next request out, honouringRetry-Afterwhere you send one. - From more than one country
- Firms serve different licensed entities by location, so the same page may be read from more than one country. robots.txt is read from each of them, in case it differs.
- Some pages are rendered
- Where the disclosure is only drawn by script, the page is opened in a headless browser that sends the same user agent. It reads; it does not click, sign in or submit anything.
Two lines of robots.txt.
The scanner matches its own name. This blocks it from the whole site and leaves every other agent's rules as they are.
User-agent: FigureCheckBot Disallow: /
A robots.txt that cannot be fetched at all is treated as permission, because the pages are ones a client has asked to have checked. An explicit rule is always obeyed. To talk to a person instead, write to hello@mvrk.systems with the host and roughly when the requests arrived.