Site Crawl
The full crawl of your site, how it budgets and samples pages, how it reports a block, and how to give it access to staging.
The site check watches a few important pages every day. The site crawl finds what nobody is looking at. It walks your site by following links from the homepage, reads every page it reaches, and reports problems across the whole site: broken pages and links, redirect chains, canonical and noindex problems, duplicate titles and content, orphan pages, insecure pages, hreflang and structured data errors.
When it runs
- Once a week for every workspace with an active Search Console connection.
- About five minutes after a production deploy is recorded.
- When you click Run Crawl on the Crawl page, or approve a crawl the agent proposed.
Only one crawl runs at a time per workspace. Crawling is metered per page fetched and paid from the account's balance, so a crawl does not start when the balance cannot cover its page budget, and it stops early if the balance runs out partway.
How it crawls
The crawl starts at https:// plus your website's host and follows links on the same site, including between the bare host and www. It never follows links to other sites. It obeys robots.txt for RankDebugBot, skips links marked rel="nofollow", and follows redirects one hop at a time so every hop is recorded. The sitemap is read first but its URLs are visited last, so a page that only the sitemap knows about shows up as an orphan instead of using up the budget.
It is deliberately gentle: a few requests in flight, at least a quarter of a second apart, fewer when responses slow down or fail. RankDebugBot describes how it identifies itself and how to allow or block it.
Page budget and sampling
A crawl fetches at most 500 pages, unless your workspace has agreed a larger budget with us. Redirect hops are recorded but do not count against the budget.
Large sites have far more pages than that, so the crawl samples by URL pattern. Pages that share a template, such as /listing/{id} for every listing, count as one pattern, and once a pattern has reached its sample size the remaining URLs of that pattern are counted but not fetched. The sample size is a twentieth of the budget, and never fewer than ten pages, so a 500-page crawl reads up to 25 pages of each pattern. This spends the budget across the whole site rather than on one long list.
A crawl also stops after two hours. The Crawl page says whether the crawl reached every page it found or stopped at its budget.
Blocked, not broken
A 401, 403 or 202 answer, a bot challenge page, or a CDN header saying the request was mitigated means the crawler was refused. RankDebug treats that as a statement about the crawler, never as an error on your page. If the homepage refuses the crawler, the crawl stops there. If a fifth or more of the fetches were refused, the crawl reports Crawler Blocked once instead of listing hundreds of false errors.
A 429, or a 503 with Retry-After, means the site asked the crawler to slow down. The crawl pauses, honours Retry-After, and after repeated rate limiting stops and reports Crawler Rate Limited.
Client-side rendered pages
Some pages arrive as an almost empty HTML shell that JavaScript fills in after it loads: very few words and links, and a lot of script. Checks such as a missing title or thin content would be false on those shells, so they are skipped there, and the crawl reports Client-Side Rendered once with the count.
To check those pages as a browser sees them, turn on Render JavaScript Pages in Crawler Access. After each crawl, a limited number of the flagged pages are then loaded in a headless browser, with the same user agent and signature, and the skipped checks run on the rendered page.
Crawler access for staging and protected sites
Crawler Access on the Crawl page sets how the crawler reaches your site. Only a workspace owner or admin can open it.
- Username and Password send HTTP basic authentication, for a staging site behind a login.
- Custom Headers send up to ten headers with every request, for example a header your bot protection lets through. A value can be up to 1,024 characters.
Host,Content-Length,User-Agent,AcceptandAuthorizationare set by the crawler and cannot be overridden. - Render JavaScript Pages turns rendering on or off, as above.
The login and header values are stored encrypted and are only loaded when a crawl runs. Once saved they are never shown again, only the header names and whether a login is set. To change the headers, type every value again.
Reading the results
The Crawl page shows the latest crawl with its progress while it runs, then:
- The findings, each with its severity, the number of pages it names and the Search Console clicks those pages earned over the last 28 days, ordered by severity and then by clicks.
- Since Last Crawl, the findings and pages that are new and those that are Resolved compared with the crawl before.
- AI Readiness, described on its own page.
Open a finding to see every page it names, with a detail per page such as the status code, the redirect target or the duplicate title, and the clicks on that page. Export CSV downloads the list with the URL, the detail and the 28-day clicks per page. A finding keeps up to 1,000 pages; the count always shows the true total.
When a crawl finds new high-severity findings compared with the crawl before, the workspace is alerted the same way as for a site check, and subscribed webhooks receive site_crawl.regression.
Every finding kind, with what it means and how to fix it, is in the findings reference.
Related documentation
- Guide
How the parts of RankDebug fit together, and where to find each one in the dashboard.
- Connections
How data sources connect to a workspace, how often they sync, and what happens when you disconnect one.
- Search Console
Connect Google Search Console with read-only access. What RankDebug pulls and what it is used for.
- Google Analytics
Connect Google Analytics with read-only access to see what search visitors did after they landed.
- Cloudflare
Connect Cloudflare with a read-only API token to see which crawlers reach your site and what the edge did to them.
- Bing Webmaster Tools
Connect Bing Webmaster Tools with an API key to add Bing search data and crawl statistics.
- GitHub
Connect a GitHub repository so production deploys are recorded on their own, with the SEO-relevant files each one changed.
- GitLab
Connect a GitLab project with a read-only token so production deploys are recorded on their own.