Site Crawls

Read crawls of the workspace website, their findings, and every page of one finding.

These endpoints need the search scope. Starting a crawl and changing crawler access are done in the dashboard. The Site Crawl guide explains how crawls run.

A crawl

Every crawl in a response has this shape:

FieldNotes
idThe crawl id
triggerdeploy, scheduled or manual
statusqueued, running, done or failed
pageCapThe page budget
pagesCrawled, pagesDiscoveredPages fetched, and URLs found
pagesSampledOutURLs skipped because their URL pattern reached its sample size
completetrue when the crawl reached every page it found within the budget
blockedPages, rateLimitedHow many fetches were refused, and whether the site rate-limited the crawl
clientRenderedPages, renderedPagesPages that arrived as client-side shells, and how many were rendered in a browser
sitemapUrls, edgeCountURLs in the sitemaps, and internal links found
aiReadinessPer AI crawler, robots.txt and edge access, and how much content is in the HTML. See AI readiness
errorWhy a crawl failed or stopped early
createdAt, startedAt, finishedAtTimestamps

A finding in a crawl report has kind, severity, count (pages it names in total), urls (the first pages), details (a short detail per URL) and clicks (Search Console clicks on those pages over the last 28 days).

Latest crawl

GET /v1/site-crawls/latest?workspace_id={id}

latest is the newest crawl in any state. crawl is the newest finished one, with its findings, and baseline is the finished crawl before it. delta lists the findings added and resolved between the two, or is null when there is nothing to compare.

List crawls

GET /v1/site-crawls?workspace_id={id}

Crawl history, newest first. limit is 1 to 100, default 20.

Get a crawl

GET /v1/site-crawls/{id}

One crawl with its findings and the delta against the crawl before it, in the same shape as the latest crawl.

Every page of one finding

GET /v1/site-crawls/{id}/findings/{kind}

kind is a crawl finding kind such as broken_page; the findings reference lists them all. The response has the finding's kind, severity, count and clicks, and urls with a url, detail and clicks for each page.

Download a finding as CSV

GET /v1/site-crawls/{id}/findings/{kind}/urls.csv

The same pages as a CSV file, with the URL, the detail and the 28-day clicks per page.

Related documentation
  • Overview

    Base URL, authentication, scopes, errors and rate limits for the RankDebug API.

  • API Keys and Scopes

    Create API keys, limit them with scopes and a workspace binding, and keep them safe.

  • Workspaces

    List, create, update and delete workspaces, and read a workspace's request logs and activity.

  • Search

    Search Console performance, top rows, opportunities and sitemaps for a workspace.

  • Traffic

    Google Analytics landing pages, acquisition and measurement health, and CDN crawler traffic, error paths and firewall rules.

  • Deploys

    Record deploys from CI, list them, and read one deploy with its site check, crawl and search impact.

  • Digests

    Read what the latest scheduled digest email for a workspace contained.

  • Reports

    Read the reports shared with a workspace, their comment threads, and their PDF.

Was this helpful?

On this page