What we check before trusting a crawl

A crawl can be quietly wrong. Soft 404s, partial renders, blocked assets, staging leakage. The unglamorous checks we run before we believe the table.

Silkra team5 min read

Summarize

A crawl looks finished when the progress bar stops. That is the dangerous moment. The table can be complete and still be wrong in ways that do not show up as an error. Soft 404s sit on 200. Partial renders look like thin pages. Staging URLs sneak into a production project. You build the whole audit on that, then spend a day defending a finding that was never on the live site.

We have trusted a crawl too early. More than once. These are the checks we run now before we let the table talk.

Finished is not the same as trustworthy

A crawler reports what it got. Status codes, extracted text, links, screenshots if you asked for them. "Got" is doing a lot of work in that sentence.

It got a response. It may not have got the page a person sees. It may have got a login wall, a cookie banner occupying the first screen, a client-rendered shell with the article still in a JSON blob, or a soft 404 that says "not found" in the body and 200 in the header.

Agents make this worse, not better. A model will write fluently from a bad table. The draft will have URLs. The URLs will be real. The claims will be about a version of the page that never quite loaded.

So the first job is not clustering or chunking. It is deciding whether the crawl is a picture of the site.

In plain terms

We do not start the interesting work until we have caught the crawl lying in a boring way.

The boring checklist

None of this is clever. That is why it works.

  • Count vs sitemap vs search console.

    If the crawl is 40% smaller than the sitemap, something was blocked, paginated away, or never linked. If it is much larger, you may have parameters, faceted traps, or a staging host riding along.

  • Status mix.

    A wall of 200s feels healthy. It can hide soft 404s, empty templates, and "success" pages that are errors in the body.

  • Indexability on pages you care about.

    Canonicalized, noindexed, or robots-blocked URLs still get crawled sometimes. They do not get to drive the story you tell a client about what Google or a model can use.

  • Rendered vs raw, on a sample.

    Open ten important URLs. Compare the extracted text to the live page. If the H1, the price, or the main explanation is missing from the extract, the rest of the table is a rumor.

  • Blocked assets.

    A page that lost its JavaScript bundle is a different page. So is one that lost the CSS that unhides the content. The crawler may still store a 200.

  • Host leakage.

    staging., www vs apex, preview deployments, localization prefixes you did not mean to include. One leaked host can dominate a "thin content" report.

We still skip a row. Then we find it in the deliverable.

Notebook sketch of a tidy table with one steel check mark in the margin, and a small question mark next to a row.

The table looks done. The check is the part that takes fifteen minutes and saves a day.

Soft 404s and empty successes

The soft 404 is the classic. The server is proud. The body says the product cannot be found, or the article moved, or the search had no results. Extraction still gets a title and a paragraph of consolation copy. A thin-content rule will flag it, or worse, will not, because the template is wordy.

We look for repeated titles, repeated H1s, and bodies that collapse to the same 80 words across dozens of URLs. That pattern is usually a template saying "nothing here" in complete sentences.

Empty successes are cousins. The app rendered the chrome and never filled the article. Word count is nav plus footer. If you embed that, you get a vector of the chrome. Sibling pages huddle. The cluster looks like a topic. It is a header.

Partial renders and the pages that matter

You do not need to hand-check 8,000 URLs. You need to hand-check the ones the audit will lean on.

Commercial templates. Pricing. The docs article everyone thinks is the source of truth. A couple of blog posts that might be competing with a solution page. If those extracts are short, stale, or missing the unique paragraph, stop. Fix the render, the wait, the blocked script, or the extractor — then crawl again.

We have published findings from a first-pass extract that lost the main column. The "thin" page was not thin. The crawler left too early. That is an expensive way to learn a patience setting.

Rate limits produce a related mess: the first 200 pages look perfect, the rest look like errors or repeats. An average across the crawl then describes the beginning of the site, not the site.

When the number looks too clean

This is the check that is hardest to teach, because it is a feeling with a method behind it.

If every important page is 200, indexable, fully rendered, and uniquely titled, maybe you did a good job scoping. Or maybe you crawled the marketing site and missed the app, the help center, or the filtered catalog where the actual problems live.

If a semantic map shows one giant cluster, maybe the brand voice is that strong. Or maybe every extract includes the same 400 words of nav.

Silkra is a crawler, so this is the unglamorous part of the product too. We would rather a workspace show you the missing render than a pretty cluster built on chrome. The cluster can wait.

We still get burned

The checklist is not a guarantee. A site will find a new way to look finished.

The question we ask before a report goes out is smaller than "is this crawl complete?" It is "which finding would fall apart if the extract is wrong?" Those findings get a second open of the live URL. If they survive, the table earned a little more trust. If they do not, we fix the crawl first and pretend the draft never happened.

Put crawl evidence to work

Download Silkra and turn audits, briefs, and fixes into one focused workflow.

Get started

Create your first workspace.

Crawl a site, then ask what needs attention.

Free to start. No credit card needed.