Site CrawlFeature

A crawler that tells you what changed, not just what is wrong

A crawl report with thousands of rows is hard to act on. Site Crawl compares every crawl with the one before it, so you can see which issues are new, which ones a deploy introduced, and which pages they affect.

Dozens of Checks

Status codes, redirects, canonicals, indexability, titles, descriptions, headings, thin and duplicate content, hreflang, structured data and security headers.

Diffed Against the Last Crawl

New high-severity findings are flagged as regressions and can alert the workspace, so a fresh problem never hides among old ones.

Built for Templates

URLs are grouped by template and sampled, so a site with millions of listings gets a useful picture of every page type without crawling everything every week.

What It Finds

The issues that keep pages out of the index

The crawler starts at your homepage, follows internal links breadth first and uses your sitemaps to find pages that nothing links to. Each page is checked as a search engine would see it, and the link graph is kept so depth and orphan pages can be measured.

  • Broken pages, broken internal links and server errors
  • Redirect chains and loops
  • Canonical conflicts, noindex pages and canonicalised pages
  • Orphan pages and pages buried deep in the site
  • Missing, duplicate or badly sized titles and meta descriptions
  • Thin or duplicate content, missing alt text, invalid hreflang and structured data errors
How It Behaves

Polite by default, thorough when you need it

The crawler identifies itself honestly as RankDebugBot, respects robots.txt for its own name, keeps concurrency low and backs off when a server asks it to. When a page is a client-side shell, you can turn on rendering so the checks run on the page as a browser builds it.

  • Weekly scheduled crawl, plus a crawl after each recorded deploy
  • Optional JavaScript rendering for client-rendered pages
  • Custom headers or basic auth for staging or bot-protected sites, stored encrypted
  • Crawl size set per contract, from a few hundred pages to a large sample of a big site

Site Crawl: questions and answers

Will the crawler slow my site down?

It is designed not to. It runs a small number of requests at a time with a pause between them, retries gently and honours Retry-After when your server asks it to slow down.

Can it crawl a site behind a firewall or bot protection?

Yes. You can give it custom headers or basic auth credentials, which are stored encrypted. Many teams also allow the crawler by its user agent in their CDN rules.

Does it render JavaScript?

When you turn rendering on, pages that arrive as an empty shell are rendered in a headless browser and checked again on the rendered page. It is off by default because most search-critical content should be in the HTML.

How many pages can it crawl?

The crawl size is part of your contract. Large sites are sampled by template, so each page type is covered without crawling every listing every week.

See Site Crawl on your own site

In a short call we connect your site, show you what RankDebug finds, and agree a plan that fits it.

Book a Demo