silkra

Crawl settings

Every control on Settings → Spider, and when to touch it.

Summarize

Open Settings with ⌘J and stay on Spider. Changes save as you go and apply to the next crawl in this workspace. Scope is not in this panel. That is set by what you paste on the start screen, then refined in preview.

How to start a crawl →

Account, theme, privacy, and connectors are other tabs. They are not crawl settings.

Rendering

JavaScript rendering: the setting that changes results.

ModeWhat it does
Auto-detectRenders a page only when it looks like it needs a browser. The default.
Always renderEvery page in a headless browser. Slower. Use it on SPAs and client-routed apps.
Static HTML onlyFetches HTML. No JavaScript. Use it when the HTML already is the page.

Limits

  • Max pages: how many pages one run can add. The same control lives in Crawl options on the start-screen preview. Your plan cap wins if you type higher.
  • Request timeout: how long Silkra waits on one page before marking it timed out. Raise it on a slow host. Leave it if pages are failing for other reasons.

Etiquette

Respect robots.txt: stays on for client sites you do not run. It honors disallow rules and crawl-delay. Turn it off only on a host you control.

Crawl pace: how hard the spider leans on the host.

PaceUse when
FastA site you own
MediumThe default. Fine for most client sites
ModerateYou want more space between requests
ConservativeThe host already rate-limits, or you do not want to lean on it
Custom overrideYou set delay and concurrency yourself

Custom reveals two more fields.

  • Request delay: the gap between request starts, in milliseconds.
  • Concurrency: how many pages run at once. Higher concurrency can get you blocked.
  • Eco mode: not in this panel. It sits in Crawl options on the start screen. A one-run slower pace so a large scrape stays cool. It does not rewrite the pace you saved here.

URL rules

  • Strip tracking params: drops utm, gclid, and the usual analytics junk so you do not crawl the same page twice. Leave it on.
  • Max redirects: stops a hop chain. The page is marked after that many hops.
  • Include only: a glob list. Empty means no extra restriction. A value like /blog/** keeps the spider inside those paths.
  • Exclude: skips matching URLs. Use it for /admin, file types, or utility paths you do not want in the table.

Folder picks on the start-screen preview write the same include and exclude lists. Patterns you add here persist for the next crawl in this workspace.

Identity

  • User-Agent: Default or Chrome, or a string you type. Leave the field empty for Default. That identifies Silkra as itself. Chrome is the one to try if the host is picky about bots. Do not impersonate another crawler.
  • Max response size: skips huge responses, in MB. Set it to 0 if you want no cap.

Reset

Reset at the top of the panel puts rendering, limits, etiquette, URL rules, and identity back to defaults. Theme and account stay put.

Refreshing a crawl → Connections →