What meaning actually means to a model
To a model, meaning is a direction in vector space, not comprehension. Why that shortcut works, and where a person and a model still disagree.
Michael Davis/5 min read
People say models understand meaning. The honest version is smaller. An embedding model places text in a direction. Pages that point the same way get called similar. We read "about the same thing" into that. The model did not sit with the page. It ran a calculation.
That shortcut is useful. It is also why a privacy policy and a product page can look close, and why a person looking at both will not.
Meaning is a direction, not a reading
A vector is a list of numbers that points somewhere. The model that produced it was trained on a lot of text: which words show up near which other words, in which neighborhoods, across a huge pile of documents.
So "meaning," in this stack, is a placement. Two passages point a similar way if the training data treated their neighborhoods as related. Cosine similarity is the number for how aligned those directions are.
That is not a person finishing the page and saying what it was about. It is not a topic tag someone chose. It is not intent. Those are stories we attach after the arrow shows up.
We say "what a page means" in this series because the cluster view is easier to use that way. The sentence is a shortcut. Keep it, and remember it is a shortcut.
In plain terms
The model placed the text. We are the ones calling the placement meaning.
Where the direction comes from
The placement comes from co-occurrence, plus whatever text you fed the model on this run.
Co-occurrence is a dry word for a simple pattern. "Starter plan" and "basic tier" hung around the same kind of sentences in training, so they often point the same way even when they share no keyword. "Pizza recipes" and "privacy policy" did not.
The text you feed still matters. Title only makes a jittery little arrow. Main content is usually the one you want. The whole render, nav and footer included, pulls sibling pages together because they share chrome. Same URL, three possible meanings, depending on the extractor.
A heading-aware chunk can point one way. A leftover sentence from the previous section, glued onto the next window, can point another. The page did not change. The object being placed did.
When a person and a model disagree
On one 105-page software crawl we already wrote about, randomly paired pages averaged about 0.82. A privacy policy and a product page could reach 0.85. No SEO would call those topically close. The model still put them nearer than the word "unrelated" wants you to feel.
That pair is the whole argument. A human reading those two URLs sees different jobs: legal text, and a thing you can buy. The model sees a lot of shared brand language, a shared template, maybe a shared product name in the footer. Close, as a direction. Not close, as a job.
The reverse happens too. Two pages can do the same job in different clothes — a docs article and a sales page, a changelog and a migration guide — and sit farther apart than you expected, because the words around the job never overlapped in training the way they overlap in your head.
Neither miss is a bug you file. Both are what you get when meaning is a neighborhood in training data.

Same direction. Different job. That gap is the part worth staying curious about.
What that does in an audit
Once you stop asking the model to "understand," the cluster view gets easier to use.
A tight group of pages is a hint they will compete for the same question, or that they are a healthy topic set. You still open them. Semantic analysis is that grouping, not a verdict.
A page sitting by itself might be an orphan in meaning-space. It might also be a legal page, a login, or a one-off tool. The arrow does not know.
A query that surfaces a changelog instead of the pricing section is not the model being cute. The changelog passage pointed nearer than anything else you stored. That can be a writing problem, a chunk problem, or a missing page. The score will not tell you which.
The useful habit is boring. Treat "similar" as "worth opening." Treat "meaning" as "this text, this model, this extraction." Then decide.
Silkra's semantic map is one way of seeing those neighbors. We built it because a spreadsheet of directions is miserable to read. The map is still a placement view, not a reading of the site.
The part that is still fuzzy
We have not found a clean way to say this in a client deck without either underselling the tool or overselling the model. "It understands the site" is the line that books the meeting. "It placed the text" is the line that keeps you from merging the wrong two URLs.
Both can be true enough to work with, if you keep the second one in your pocket.
What we have not measured yet is how often the model's closest pair is a pair a senior SEO would also pick, versus how often it is the privacy-and-product kind of neighbor. One crawl gave us the example. It did not give us the rate. That is the next count worth making — on a different site, so the 105-page software crawl does not have to carry every post.
Put crawl evidence to work
Download Silkra and turn audits, briefs, and fixes into one focused workflow.
Get started›