Embeddings vs keywords: what each can't see
Keywords see demand and exact language. Embeddings see topical neighbors and miss volume. What each approach gets wrong, and why the job still uses both.
Silkra team/5 min read
Do you still need keyword research if pages are vectors? Yes. They are different sensors. Keywords see demand and exact language. Embeddings see topical neighbors and miss volume entirely.
The interesting part is not the debate. It is the failure cases. Each one goes blind in a place the other can still see.
They are looking at different jobs
Keyword research asks what people type, how often, and in what form. It is a demand map, plus a language map. "starter plan pricing" and "basic tier cost" can be two rows with two volumes even when a person would call them the same question.
Embeddings ask which texts point the same way. Those two phrases often sit close, which is why semantic analysis can group pages that never shared a primary keyword.
One sensor is about the market. The other is about the text you already have. Mixing them up is how a cluster view gets treated like a ranking factor, and how a keyword export gets treated like a picture of what a page is about.
In plain terms
Keywords tell you what people asked for. Embeddings tell you which of your pages are hanging around the same idea.
What keywords cannot see
Keywords miss the paraphrase. If the only row you have is "enterprise SSO," a page titled "how single sign-on works on the business plan" can look like a miss in the export and a hit to a buyer.
They also miss structure. A site can target the right terms and still have three URLs doing the same job, or a docs page carrying the explanation the commercial page never makes. The keyword map will not show that overlap. A neighbor list will.
And they miss retrieval. Answer engines do not start by counting query matches the way a 2014 ranking factor slide described. They embed a question, embed chunks, and pull the nearest passages. A page can "own" the keyword in your tracker and still not be the passage that gets used.
None of that retires keyword research. It bounds it. Demand is still demand.
What embeddings cannot see
Embeddings do not see volume. Close is not popular. A tight cluster of three forgotten blog posts is still a tight cluster.
They are weak on the words you care about most as an SEO: brand names, product names, SKUs, versions, and numbers. Those tokens are sparse, or they get smoothed into a neighborhood with every other product name the model has seen. "Plan 2" and "Plan 3" can sit closer than the pricing difference deserves.
They also cannot tell you that two pages serve different intents. A comparison page and a docs page can point the same way and still be the wrong merge.
On the 105-page software crawl we already published numbers for, a privacy policy and a product page could reach 0.85. Cosine similarity was doing its job. Keyword research would not have grouped those two. In that case the keyword view was the one that matched a human.

Two instruments. Different blind spots. The drawing is the whole post if you want it that short.
Where each one gets a page wrong
A handful of patterns show up enough to keep on a scrap of paper. These are teaching cases, not a study.
| What you asked | Keywords tend to miss | Embeddings tend to miss |
|---|---|---|
| Same job, different words | The paraphrase | — |
| How many people want it | — | All of it |
| Brand / SKU / version | Less often | The exact string |
| Two pages a person would not merge | Less often | Shared voice and template |
| Which passage a question will hit | The URL-level row | A bad chunk still wins if it is nearest |
The blank cells are the point. You are not choosing a winner. You are noticing which sensor is allowed to speak.
Brand terms are the one we keep tripping on. A client will ask why the model cannot "see" the product name. It can see a direction that includes a lot of product-shaped text. That is not the same as matching the string.
You still use both
In the actual job it looks like this.
Start with the questions people ask, and the language they use. That is still keyword work, plus sales calls, plus search console, plus whatever else you already trusted.
Then look at whether the site has a passage that sits near those questions, and whether two pages are sitting on top of each other. That is embedding work.
If the keyword exists and no nearby passage does, you have a gap. If the passage exists and no one searches that way, you have a writing choice, not a demand proof. If both sensors agree a cluster is crowded, you probably do have a merge or an internal-link problem.
Silkra is built around the second sensor. We embed a crawl so we can see neighbors. We do not pretend that view replaced the first sensor. Volume never showed up in a vector.
The part that is still fuzzy
The industry line is that embeddings replaced keywords. The quieter line is that they replaced overlap metrics for comparing text, and left demand research where it was.
What we have not done yet is sit with one client question set and score, side by side, which pages a keyword export would have sent us to versus which passages a nearest-arrow search actually pulled. That bake-off is the post behind this one. Until we run it, the table above is a field sketch: useful, and waiting on a count.
Put crawl evidence to work
Download Silkra and turn audits, briefs, and fixes into one focused workflow.
Get started›