RankDebugBot
The crawler RankDebug uses to audit websites for their owners. How to recognise it, how it behaves, and how to allow or block it.
RankDebugBot is the crawler behind RankDebug. It fetches a website on behalf of that website's owner, to find what stops the site's pages from being crawled, indexed and ranked: broken pages, redirect chains, blocked crawlers, missing or conflicting canonicals, noindex tags, duplicate titles and similar problems. It does not train AI models, resell content or crawl the web at large.
How to recognise it
Every request carries this user agent:
RankDebugBot/1.0 (+https://rankdebug.com)Requests are also signed with an HTTP message signature (the Web Bot Auth scheme), so a CDN can verify that a request really comes from us rather than from someone copying the user agent. Each signed request carries Signature-Agent, Signature-Input and Signature headers, and the public key is published at:
https://api.rankdebug.com/.well-known/http-message-signatures-directoryWhich sites it visits
RankDebugBot only visits a website that someone has added to their RankDebug workspace. It does not follow links off that site.
- Site crawl. Walks the site's own pages by following links, starting at the homepage, with the sitemap read afterwards. It runs once a week for sites whose Search Console property is connected, when the owner starts one, and a few minutes after the owner reports a deploy.
- Site check. Fetches robots.txt, the sitemaps and at most 30 pages: the homepage and the pages that earn the site the most search traffic. It runs once a day and after a reported deploy, to catch a release that broke something important.
- Page rendering. When the owner turns it on, a small number of pages whose content only appears after JavaScript runs are loaded in a headless browser, with the same user agent and signature.
How it behaves
- Follows robots.txt. The site crawl obeys the rules for
RankDebugBot, or for*when no group names it. Links markedrel="nofollow"are not followed. - Crawls slowly. At most six requests in flight to a site, at least a quarter of a second apart, and fewer as soon as responses slow down or fail.
- Backs off when asked. A
429or503pauses the crawl, honouringRetry-Afterwhen the site sends one. After repeated rate limiting the crawl stops for the day and reports it. - Stops at a page budget. A crawl fetches at most 500 pages unless the site's owner has agreed a larger budget with us. Large sites are sampled: a few pages of each URL pattern, never every listing.
- Reads, never writes. Only
GETrequests. It never submits forms, logs in, or changes anything, unless the site's owner gave it a staging login or header for their own site.
How to control it
To keep RankDebugBot out of the whole site:
User-agent: RankDebugBot
Disallow: /To keep it out of one section:
User-agent: RankDebugBot
Disallow: /search/If you own the site and a firewall or bot protection is blocking it, allow requests whose user agent contains RankDebugBot, or rely on the signature where your CDN supports Web Bot Auth. RankDebug reports a blocked crawl as "we were blocked" rather than as errors on your pages, so a block never shows up as false problems.
Questions or a problem
If RankDebugBot is causing trouble on a site you run, use the contact form and include the site and the time. We reply to crawler reports first.