FigureCheck™ · the scanner

FigureCheckBot

If this address appeared in your logs, a FigureCheck™ client has asked us to check the risk disclosure on a page of yours, usually because the page promotes their brand. This is what the scanner does, what it does not do, and how to stop it.

User-Agent  FigureCheckBot/0.1 (+https://mvrk.systems/figurecheck/bot)
01 · What it reads

One page at a time,
and only pages it was given.

It does not crawl. Every request is for an address a client named, or a partner page they declared, and it reads that page for one thing: the retail-loss percentage in the disclosure.

robots.txt first
Fetched before the page and honoured, including Crawl-delay. An explicit Disallow is always obeyed; the page is then reported to the client as unreachable rather than read some other way.
One request per host
At most one request every two seconds to any one host, or slower if your Crawl-delay asks for it. The limit is per host, so no single site absorbs the scanner's volume.
Backs off when asked
A 429 or a 5xx pushes the next request out, honouring Retry-After where you send one.
From more than one country
Firms serve different licensed entities by location, so the same page may be read from more than one country. robots.txt is read from each of them, in case it differs.
Some pages are rendered
Where the disclosure is only drawn by script, the page is opened in a headless browser that sends the same user agent. It reads; it does not click, sign in or submit anything.
02 · Stopping it

Two lines of robots.txt.

The scanner matches its own name. This blocks it from the whole site and leaves every other agent's rules as they are.

User-agent: FigureCheckBot
Disallow: /

A robots.txt that cannot be fetched at all is treated as permission, because the pages are ones a client has asked to have checked. An explicit rule is always obeyed. To talk to a person instead, write to hello@mvrk.systems with the host and roughly when the requests arrived.