Embeddings vs keywords: what each can't see

Keywords see demand and exact language. Embeddings see topical neighbors and miss volume. What each approach gets wrong, and why the job still uses both.

Silkra team5 min read

Summarize

Do you still need keyword research if pages are vectors? Yes. They are different sensors. Keywords see demand and exact language. Embeddings see topical neighbors and miss volume entirely.

The interesting part is not the debate. It is the failure cases. Each one goes blind in a place the other can still see.

They are looking at different jobs

Keyword research asks what people type, how often, and in what form. It is a demand map, plus a language map. "starter plan pricing" and "basic tier cost" can be two rows with two volumes even when a person would call them the same question.

Embeddings ask which texts point the same way. Those two phrases often sit close, which is why semantic analysis can group pages that never shared a primary keyword.

One sensor is about the market. The other is about the text you already have. Mixing them up is how a cluster view gets treated like a ranking factor, and how a keyword export gets treated like a picture of what a page is about.

In plain terms

Keywords tell you what people asked for. Embeddings tell you which of your pages are hanging around the same idea.

What keywords cannot see

Keywords miss the paraphrase. If the only row you have is "enterprise SSO," a page titled "how single sign-on works on the business plan" can look like a miss in the export and a hit to a buyer.

They also miss structure. A site can target the right terms and still have three URLs doing the same job, or a docs page carrying the explanation the commercial page never makes. The keyword map will not show that overlap. A neighbor list will.

And they miss retrieval. Answer engines do not start by counting query matches the way a 2014 ranking factor slide described. They embed a question, embed chunks, and pull the nearest passages. A page can "own" the keyword in your tracker and still not be the passage that gets used.

None of that retires keyword research. It bounds it. Demand is still demand.

What embeddings cannot see

Embeddings do not see volume. Close is not popular. A tight cluster of three forgotten blog posts is still a tight cluster.

They are weak on the words you care about most as an SEO: brand names, product names, SKUs, versions, and numbers. Those tokens are sparse, or they get smoothed into a neighborhood with every other product name the model has seen. "Plan 2" and "Plan 3" can sit closer than the pricing difference deserves.

They also cannot tell you that two pages serve different intents. A comparison page and a docs page can point the same way and still be the wrong merge.

On the 105-page software crawl we already published numbers for, a privacy policy and a product page could reach 0.85. Cosine similarity was doing its job. Keyword research would not have grouped those two. In that case the keyword view was the one that matched a human.

Notebook sketch of two small sensors side by side. One is labeled demand, the other neighbors.

Two instruments. Different blind spots. The drawing is the whole post if you want it that short.

Where each one gets a page wrong

A handful of patterns show up enough to keep on a scrap of paper. These are teaching cases, not a study.

What you askedKeywords tend to missEmbeddings tend to miss
Same job, different wordsThe paraphrase—
How many people want it—All of it
Brand / SKU / versionLess oftenThe exact string
Two pages a person would not mergeLess oftenShared voice and template
Which passage a question will hitThe URL-level rowA bad chunk still wins if it is nearest

The blank cells are the point. You are not choosing a winner. You are noticing which sensor is allowed to speak.

Brand terms are the one we keep tripping on. A client will ask why the model cannot "see" the product name. It can see a direction that includes a lot of product-shaped text. That is not the same as matching the string.

You still use both

In the actual job it looks like this.

Start with the questions people ask, and the language they use. That is still keyword work, plus sales calls, plus search console, plus whatever else you already trusted.

Then look at whether the site has a passage that sits near those questions, and whether two pages are sitting on top of each other. That is embedding work.

If the keyword exists and no nearby passage does, you have a gap. If the passage exists and no one searches that way, you have a writing choice, not a demand proof. If both sensors agree a cluster is crowded, you probably do have a merge or an internal-link problem.

Silkra is built around the second sensor. We embed a crawl so we can see neighbors. We do not pretend that view replaced the first sensor. Volume never showed up in a vector.

The part that is still fuzzy

The industry line is that embeddings replaced keywords. The quieter line is that they replaced overlap metrics for comparing text, and left demand research where it was.

What we have not done yet is sit with one client question set and score, side by side, which pages a keyword export would have sent us to versus which passages a nearest-arrow search actually pulled. That bake-off is the post behind this one. Until we run it, the table above is a field sketch: useful, and waiting on a count.

Put crawl evidence to work

Download Silkra and turn audits, briefs, and fixes into one focused workflow.

Get started

Create your first workspace.

Crawl a site, then ask what needs attention.

Free to start. No credit card needed.