Site Check
The daily check of robots.txt, sitemaps, hosts and your most important pages, and how it reports issues and regressions.
The site check is a quick daily look at the things that break search traffic fastest. It is narrow on purpose: rather than walking the whole site, it watches the handful of files and pages that matter most, every day and after every production deploy, so a release that breaks something important is caught within minutes.
When it runs
- Once a day for every workspace that has a website.
- About two minutes after a production deploy is recorded.
- When you click Check Site Now on the Deploys page. Only one check runs at a time.
What it checks
- robots.txt: whether it answers, its rules, and whether those rules let Googlebot, Bingbot and the common AI crawlers (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended and CCBot) reach the homepage.
- Sitemaps: up to three, either those robots.txt declares or
/sitemap.xml. A sitemap index is followed one level down so the count is pages, not child sitemaps. - Hosts: the homepage on
httpsandhttp, with and withoutwww, each followed to where it lands, to catch a site served twice. - Pages: up to 30 pages, the homepage plus the pages on your own host that earn the most Search Console clicks. Each is fetched once, and the check records its status, redirects, title, meta robots and
X-Robots-Tag, canonical, hreflang links and H1 count. - Performance: a mobile lab test of the homepage, which gives a performance score.
The check is made with the RankDebugBot user agent.
Issues and regressions
The results on the Deploys page are split in two.
Issues describe what is wrong right now, from this check alone: robots.txt unreachable, a crawler blocked, no working sitemap or an empty one, the site served on more than one host, a page that errors, a page marked noindex, a canonical that points to another host.
Regressions describe what got worse since the previous check: a page that used to answer and now errors, a new redirect, noindex added, a canonical removed or changed, a title removed or changed, hreflang links dropped, an H1 removed, a crawler that was allowed and is now blocked, a new disallow rule, a sitemap lost or shrunk by more than a fifth, the site newly served on a second host, or the homepage performance score down by 15 points or more. A regression is only reported for a page that was in both checks.
Each finding carries a severity and the Search Console clicks its page earned over the last 28 days, and the list is ordered by severity and then by clicks. Every kind is explained in the findings reference.
Alerts
When a check finds a high-severity regression, the workspace hears about it once:
- An in-app notification linking to Deploys, and an email listing the regressions, for every member who has Site Health switched on in notifications.
- A
site_check.regressionevent to any webhook subscribed to it.
Issues on their own do not alert. Only something that got worse against the previous check does, so a problem is announced once and then stays listed under issues until it is fixed.
Related documentation
- Guide
How the parts of RankDebug fit together, and where to find each one in the dashboard.
- Connections
How data sources connect to a workspace, how often they sync, and what happens when you disconnect one.
- Search Console
Connect Google Search Console with read-only access. What RankDebug pulls and what it is used for.
- Google Analytics
Connect Google Analytics with read-only access to see what search visitors did after they landed.
- Cloudflare
Connect Cloudflare with a read-only API token to see which crawlers reach your site and what the edge did to them.
- Bing Webmaster Tools
Connect Bing Webmaster Tools with an API key to add Bing search data and crawl statistics.
- GitHub
Connect a GitHub repository so production deploys are recorded on their own, with the SEO-relevant files each one changed.
- GitLab
Connect a GitLab project with a read-only token so production deploys are recorded on their own.