How to start a crawl
Paste a domain, a subfolder, or a sitemap. Preview the scope. Then scrape.
After sign-in you land on New workspace. The field accepts a domain, a subfolder path, or a sitemap URL. The paperclip uploads a CSV or TXT list. Workspaces you already have sit in the left rail.
Enter previews first. Nothing is crawled until you confirm the scope.
What to enter
Silkra picks the mode from what you type. Change the URL if the guess is wrong.
| You enter | Mode | What it crawls |
|---|---|---|
| A host, like client.com | Spider, full domain | Follows links from the start URL across that host |
| A path, like client.com/blog | Spider, subfolder | Stays under that path |
| A sitemap URL, like client.com/sitemap.xml | Sitemap | Only URLs listed in that file |
| A CSV or TXT via the paperclip | List | Only the URLs in the file |
A host like blog.client.com is a domain crawl of that host. It does not pull in shop.client.com.
Use a subfolder when the job is one slice: /blog, /en, /docs. The table stays smaller and the map stays focused. Use a sitemap when you already trust that index. Use a list for a competitor set or a handful of money pages you do not want the spider wandering past.
Preview the scope
Enter, or the arrow, maps the site before you spend the crawl. Counts come from public sitemaps and homepage links. The scrape may find more or fewer pages.
On a domain, Folders opens the map. Click a folder once to include it. Click again to exclude it. Click a third time to clear the mark. Search if the list is long. Custom adds a path Silkra did not find.
Leave every folder unmarked to scrape the full domain. Include a few and only those run. Exclude a few and everything else runs.
On a subfolder, preview counts that path. If the folder has nested paths, locales under /blog for example, you get the same folder picker scoped to those children. Unmarked means scrape the whole subfolder.
If the path you typed is not on the site, Silkra says so and shows the folders it did find. If the folder exists but is not in the sitemap, the count is unknown. You can still start. Use the page limit to keep the run bounded.
If there is no sitemap, the spider discovers pages as it goes. Target specific folders if you still want to constrain it.
If the host is unreachable, Edit URL and try again.
A sitemap index lets you pick child sitemaps. A leaf sitemap or a list shows the count and you start from there.
Page limit and Eco mode
Crawl options sits at the bottom of the preview.
- Page limit: how many pages this run will take. Your plan cap wins if you type higher.
- Eco mode: a one-run slower pace so a large scrape does not heat the machine. It does not change the pace you saved in Settings. Silkra recommends it on larger planned runs.
Rendering, robots, and include/exclude patterns live in Settings → Spider. ⌘J opens that panel.
Helper chat
A small helper sits on this screen, separate from the workspace agent. It only knows setup: what to paste, which folders to include, crawl settings. Ask it to pick folders if the map is large.
During the crawl
The table steps aside for a progress view: counts, recent URLs, and the scrape and map lanes.
- Pause: holds the run. Resume from the same control.
- Stop: keeps what you already have.
- Cancel: tears the run down.
If the host rate-limits or blocks the spider, the notice shows there. Slow the pace or switch the user-agent in settings, then try again.
A crawl that dies mid-flight leaves a banner on the table: resume, keep the partial run, or delete that version.
When it finishes
The table fills as intelligence wraps up. The map unlocks after embeddings and clusters are ready. The first intelligence run on a machine may download an on-device model, about 35MB, once.
A crawl is a version inside a workspace. ⌘N starts another run without deleting the last one. If you meant to recrawl the same site on purpose, refresh the workspace instead of opening a new one.
Workspaces → Refreshing a crawl → Crawl table → Semantic map →