Field reference
Labels and meanings for every field you can show, hide, or filter on the crawl table.
Groups match the column picker. A filter-only field works in the command bar but cannot become a column. Search Console fields appear after that connector is on.
After analysis marks fields that fill in after Silkra finishes the crawl's intelligence pass.
Badge guide
- ColumnCan be shown or hidden in the crawl table's column picker.
- Filter onlyCan narrow the crawl but cannot be displayed as a table column.
- Search ConsoleAppears after the Search Console connector is on.
- After analysisFills in after Silkra finishes post-crawl analysis.
126 of 126 fields
URL & Crawl
| Field | What it means | Available as |
|---|---|---|
Address url | The complete web address of the page | Column |
Path pathname | The URL path after the domain (e.g., "/blog/article") | Column |
Domain domain | The website domain (e.g., "example.com") | Column |
HTTPS isHttps | Whether the page uses secure HTTPS protocol | Column |
Issues issueCount | Number of issues detected on this page by the issue registry engine. | ColumnAfter analysis |
Status statusCode | HTTP response code (200=OK, 301/302=redirect, 404=not found, 500=error) | Column |
Load Time responseTime | Time to load (under 200ms=fast, over 1000ms=slow) | Column |
Redirect To redirectUrl | Destination URL after redirects. Only shown for redirected pages. | Column |
Redirect Count redirectCount | Number of redirects before reaching final URL | Column |
JS Rendered renderedWithJs | Whether this page's content came from a JavaScript render (browser) rather than the raw HTTP response. | Column |
Pre-JS Word Count sourceWordCount | Visible words in the raw HTTP response before JavaScript ran. Only measured on JS-rendered pages — compare against Word Count to see how much content depends on JS. | Column |
Discovered Via discoveredVia | How URL was discovered (spider, sitemap, manual). | Column |
Error errorType | Type of error if request failed (timeout, blocked, network). | Column |
Error Details errorMessage | Detailed error message if request failed. | Column |
Crawled crawledAt | When this page was crawled. | Column |
URL Section section | First two URL path segments. | Filter only |
URL Subfolder subfolder | First URL path segment. | Filter only |
URL Segment 1 urlSegment1 | First URL path segment without slashes. | Filter only |
URL Segment 2 urlSegment2 | Second URL path segment. | Filter only |
URL Segment 3 urlSegment3 | Third URL path segment. | Filter only |
Path Depth pathDepth | Number of URL path segments. | Filter only |
Indexability
| Field | What it means | Available as |
|---|---|---|
Indexability indexability | Status: indexable, noindex, blocked, redirected, error, or skipped (non-HTML/oversized). | Column |
Canonical canonical | The "official" URL for this content. | Column |
Indexability Reason indexabilityReason | Explains why page is not indexable (e.g., "Meta robots noindex"). | Column |
Nofollow isNofollow | Has nofollow directive - links on page pass no equity. | Column |
In Sitemap inSitemap | Whether this URL was found in the XML sitemap. Pages missing from sitemap may not be discovered by crawlers. | Column |
Page Identity
| Field | What it means | Available as |
|---|---|---|
Title title | The page title shown in browser tabs and search results. | Column |
Title Length titleLength | Character count. Aim for 50-60 characters. | Column |
Title Pixels titlePixelWidth | Pixel width in Google SERPs. Truncates around 580px. | Column |
Meta Desc metaDescription | Summary shown in search results. 150-160 characters ideal. | Column |
Meta Desc Len metaDescriptionLength | Character count. Aim for 150-160 characters. | Column |
Language language | Page language code (e.g., "en", "es"). | Column |
Meta Desc Pixels metaDescriptionPixelWidth | Pixel width in Google SERPs. Truncates around 920px. | Column |
H1 h1 | Main heading of the page. | Column |
H1 Count h1Count | Number of H1 tags. Best practice is exactly one. | Column |
First Paragraph firstParagraph | Opening paragraph text. | Column |
Content Quality
| Field | What it means | Available as |
|---|---|---|
Published Date publishedDate | When the content was first published. | Column |
Modified Date modifiedDate | When the content was last updated. | Column |
H2 h2 | Section headings. Outlines the main sections of content. | Column |
H2 Count h2Count | Number of H2 headings. | Column |
Total Words wordCount | All visible words on the page, including navigation and template text. Compare with Content Words to see how much is chrome. Under 300 = thin content. | Column |
Content Words comprehensiveContentWordCount | Words in the extracted page content — nav, footer, and other chrome removed. The useful “how much content is here” number. | Column |
Thin Content thinContent | Page has less than 300 words - may lack substance for SEO. | Column |
Unique Content % uniqueContentRatio | Share of this page that is unique to it versus nav, footer, and sidebar chrome. On-page boilerplate — not how much this URL duplicates other pages (that is Shared Passages %). | Column |
Page Structure semanticStructureScore | How clearly the page uses semantic HTML, landmarks, and heading order to describe its content. | Column |
Structure Signal semanticStructureLabel | Plain-language structure category based on semantic elements, landmarks, and heading order. | Column |
Main Content Area hasMainLandmark | Whether the page declares a primary content landmark with <main> or role="main". <article> is a grouping element, not a landmark. | Column |
Semantic Elements semanticElementCount | Count of semantic structural elements such as article, section, nav, main, aside, header, and footer. | Column |
Heading Order headingOrderOk | Whether heading levels follow a clear order without skipping levels. | Column |
Avg Sentence avgSentenceLength | Average words per sentence. | Column |
Content Age contentAgeDays | Days since content was last updated (from published/modified dates). | Column |
Content Elements contentElements | Rich elements detected in the main content area: Tables, Code Blocks, Videos. Content-scoped (excludes nav/footer). | Column |
Page Value pageValue | How important this page is relative to the rest of the site, based on inbound links and content depth. | ColumnAfter analysis |
Content Text mainContentText | Full extracted main-content text of the page. Use contains to find pages mentioning a word or phrase anywhere in their content. | Filter only |
Readability fleschReadingEase | Flesch Reading Ease (0-100). Higher = easier. 60-70 ideal for web. | Column |
AI Readiness
| Field | What it means | Available as |
|---|---|---|
Extractability contentExtractability | Legacy 0–100 summary retained for historical crawls. Prefer the specific extraction dimensions for diagnosis. | Column |
Page Shape pageContentShape | Deterministic structural classification used to choose an appropriate extraction strategy. | Column |
Page Shape Confidence pageContentShapeConfidence | Confidence in the deterministic page-shape classification (0–100). | Column |
Content Availability contentAvailability | Whether the page exposes a substantial amount of recoverable content. | Column |
Content Boundary contentBoundaryClarity | How clearly semantic markup identifies the page-specific content region. | Column |
Extraction Coverage extractionCoverage | Plain-language share of visible page content retained by the comprehensive extraction lane. | Column |
Extraction Coverage % extractionCoveragePercent | Share of visible body words retained by comprehensive extraction (0–100). | Column |
Rendering Dependence renderingDependence | Whether substantial content appeared only after JavaScript rendering. | Column |
Markup Complexity markupComplexity | How much scripts, styles, and markup surround the visible content; not a defect by itself. | Column |
Extraction Recovery extractionRecoveryLevel | Whether standard semantic extraction worked or a broader fallback was required. | Column |
First RAG Chunk firstRetrievableChunk | First ~150 words of main content — what a RAG pipeline retrieves first. | Column |
Data Tables tableCount | Number of data tables in the main content. Tables are high-value evidence for AI answer extraction. | ColumnAfter analysis |
FAQ Pairs faqPairCount | Number of question-heading -> answer pairs (FAQ / Q&A shape) in the main content. The single strongest AEO signal — pages that directly answer questions. | ColumnAfter analysis |
List Blocks listBlockCount | Number of bulleted (unordered) list blocks (2+ items) in the main content. Lists are answer-friendly structure for features and key points. | ColumnAfter analysis |
Code Samples codeBlockCount | Number of fenced code samples in the main content. | ColumnAfter analysis |
Stat Mentions statMentionCount | Number of sentences carrying a concrete statistic (%, currency, multiplier, or large number). Quotable proof points that AI answers favor. | ColumnAfter analysis |
Ordered Lists stepListCount | Number of numbered (ordered) list blocks (2+ items) in the main content — sequences, rankings, or steps. A neutral structure signal; an ordered list is NOT necessarily a how-to procedure. | ColumnAfter analysis |
Definitions definitionCount | Number of glossary-style "Term — definition" paragraphs in the main content. Strong answer-engine signal for "what is X" queries. | ColumnAfter analysis |
Pull Quotes quoteCount | Number of blockquotes / pull quotes in the main content. Often testimonials or authoritative statements worth surfacing. | ColumnAfter analysis |
Figures figureCount | Number of standalone body images (figures) in the main content. | ColumnAfter analysis |
Text vs Markup % textRatio | Visible text as a share of HTML after script and style payloads are removed. Low means the page is markup-heavy relative to what readers see — not how much JavaScript shipped. Under 10% plus thin recovered content flags Text Buried in Code. | Column |
llms.txt hasLlmsTxt | Site has /llms.txt — the emerging standard for AI crawler discovery. | Column |
AI Bot Access aiCrawlerDirectives | Per-bot robots.txt stance: which AI crawlers are allowed or blocked. | Column |
Semantic Analysis
| Field | What it means | Available as |
|---|---|---|
Semantic Group semanticGroup | Group of pages that communicate similar meaning, derived from semantic similarity. | ColumnAfter analysis |
Semantic Fit semanticClusterStatus | How confidently the page fits its semantic group: core, related, peripheral, misfit, outlier, or ungrouped. | ColumnAfter analysis |
Likely Group likelySemanticGroup | Nearest semantic group for pages that are loosely related or need topic alignment work. | ColumnAfter analysis |
Cluster Confidence semanticConfidence | Crawl-mean-centered cosine similarity to the assigned semantic group centroid. Genuine members score ~0.3-0.7; near 0 or negative means the page barely fits its group. | ColumnAfter analysis |
Closest Neighbor closestNeighbor | The most semantically similar page within this page's semantic group. Empty when no sibling clears the similarity floor ("No clear neighbor"). | ColumnAfter analysis |
Neighbor Similarity closestNeighborSimilarity | Cosine similarity between this page and its closest semantic neighbor, scored 0-1. Higher means the two pages cover closer meaning. | ColumnAfter analysis |
Key Concepts keyConcepts | Core topics extracted from the page content — natural phrases like "content marketing" or "semantic search". Broader than named entities. | ColumnAfter analysis |
Cannibalization cannibalizationRisk | Cannibalization risk: near-identical page vector to another page in the same semantic group (raw cosine ≥ 0.96). Diffuse/Ungrouped groups are excluded; passage-level clones use Shared Passages % instead. | ColumnAfter analysis |
Intro Matches Body firstChunkPageAlignment | Centered cosine between the opening passage and the mean of the rest of the page. Lower means the intro may promise something the body does not deliver. Blank when the page has only one passage. | ColumnAfter analysis |
Title Matches Body titleContentAlignment | Cosine similarity (0-100) between the page title's embedding and the page's content embedding. Low values suggest the title promises something the body does not deliver. | ColumnAfter analysis |
Passage Agreement topicFocus | Average centered cosine of each content passage to the page’s passage centroid — how much the passages agree with each other. Low values mean the page drifts across topics or buries an unrelated section (see Weakest Section). Blank on single-passage pages. | ColumnAfter analysis |
Best Passage bestPassage | The passage most representative of what this page is about — the section a retrieval system would most likely quote. The position prefix (e.g. 5/8) shows where it sits: a best passage deep in the page is a buried lead worth moving up. | ColumnAfter analysis |
Weakest Section weakestSection | The passage least aligned with the rest of the page — the section dragging Passage Agreement down. A candidate to split into its own page, tighten, or remove. | ColumnAfter analysis |
Passages passageCount | How many retrievable passages the page splits into — the chunks an LLM retrieval pipeline would index. Boilerplate and near-empty sections are excluded. | ColumnAfter analysis |
Shared Passages % sharedContentShare | Percent of this page’s passages that have a near-duplicate on at least one other page (chrome already stripped). 70+ is usually a clone or doorway set. Mid-range is often shared template blocks. Not Unique Content % — that one is chrome on this page. | ColumnAfter analysis |
Repeated Passage % repeatedPassageShare | Share of this page’s passages that are repeated across 5+ other pages — the “every post ends with the same pitch” pattern. Distinct from Shared Passages %, which also counts a match on just one other page. | ColumnAfter analysis |
Shared With sharedContentPages | How many other pages share at least one near-identical passage with this page. Combined with Shared Passages %, this identifies duplicate families: many partners at a high share means a template-generated set. | ColumnAfter analysis |
Site Structure
| Field | What it means | Available as |
|---|---|---|
Depth depth | Clicks from start URL (lower = easier to find) | Column |
In-Group Linking linkStatus | How well this page is internally linked to other pages in its own semantic group (Well-connected → Moderate → Weak → Isolated). A linking signal scoped by topic — distinct from site-wide structuralRole. | ColumnAfter analysis |
Links (In/Out) linkCount | Combined internal and external link counts displayed together. | ColumnAfter analysis |
Click Depth clickDepth | Fewest clicks from the homepage to reach this page, following links anywhere on the page (navigation, body, or footer). | ColumnAfter analysis |
Internal uniqueInternalLinkCount | Unique links to same domain. Important for site structure. | Column |
Internal Importance pageRankScore | How important this page is within the whole-site internal link structure — counting every internal link (navigation, body, and footer). Higher means more of the site points here. | ColumnAfter analysis |
External uniqueExternalLinkCount | Unique links to other domains. Outbound link equity. | Column |
Structural Role structuralRole | This page's role in the whole-site link structure: hub (links out to many pages), bridge (connects different topic areas), spoke (normal page), or orphan (no internal links point to it). | ColumnAfter analysis |
Inbound Internal inboundInternalCount | How many internal pages link to this page across the whole site (navigation, body, and footer links). | ColumnAfter analysis |
Nofollow Links nofollowLinkCount | Links with rel="nofollow". | Column |
Dead End isDeadEnd | Page has no outbound internal links. | Column |
Empty Anchors hasEmptyAnchors | Has links with no anchor text. | Column |
In-Body Links inBodyLinkCount | Links parsed from main-content markdown (editorial body links). Replaces the old DOM-boilerplate content-link heuristic. | Column |
In-Body Internal inBodyInternalLinkCount | In-body links pointing to the same domain. | Column |
In-Body External inBodyExternalLinkCount | In-body links pointing to other domains. | Column |
Internal Link URLs internalLinkUrls | URLs of all outgoing internal links on the page (nav + body). Use contains to find pages linking to a URL or path, e.g. "/pricing". | Filter only |
External Link URLs externalLinkUrls | URLs of all outgoing external links on the page. Use contains to find pages linking out to a domain, e.g. "facebook.com". | Filter only |
Anchor Texts anchorTexts | Anchor text of every outgoing link on the page (internal + external). Use contains to find pages with links labeled a certain way, e.g. "learn more". | Filter only |
Media & Assets
| Field | What it means | Available as |
|---|---|---|
Images imageCount | Total number of images on page. | Column |
Missing Alt imagesMissingAlt | Images without alt attribute. | Column |
Image Details images | All image objects with src, alt, dimensions. | Column |
Structured Data
| Field | What it means | Available as |
|---|---|---|
Schema Types schemaTypes | All JSON-LD @type values found (Article, Product, FAQPage, etc.). | Column |
Schema Count schemaCount | Number of JSON-LD script blocks. Filter: > 0 = has schema. | Column |
Social Preview
| Field | What it means | Available as |
|---|---|---|
OG Title ogTitle | Open Graph title (og:title). | Column |
OG Description ogDescription | Open Graph description (og:description). | Column |
OG Image ogImage | Open Graph image URL (og:image). | Column |
Internationalization
| Field | What it means | Available as |
|---|---|---|
Hreflang Tags hreflang | All hreflang entries (language/region alternates). | Column |
Search Console
| Field | What it means | Available as |
|---|---|---|
Clicks gsc_clicks | Clicks from Google Search over the last 28 days | Search Console |
Impressions gsc_impressions | Impressions in Google Search over the last 28 days | Search Console |
CTR gsc_ctr | Click-through rate from search results over the last 28 days | Search Console |
Avg Position gsc_position | Average ranking position in Google Search over the last 28 days | Search Console |