What is a vector in SEO?

A vector is a list of numbers that points in a direction. Why SEO tools started using them, and what close in vector space can and cannot tell you.

Michael Davis8 min read

Summarize

People keep saying "vector" in semantic SEO like everyone already has one in their pocket. A vector is a list of numbers that points in a direction. That's it. When a tool says two pages are close in vector space, it means those two lists ended up pointing a similar way.

The list is long. Hundreds of numbers, sometimes a couple thousand. Each one is a coordinate on an axis the model invented. In a spreadsheet those columns would have names. Here they don't. You can't open the list, find the 47th number, and say "that's the pricing score." The useful question is simpler: which way is the whole list pointing, compared with another list?

What a vector is

If you remember anything from geometry, a vector is an arrow. Length plus direction. On a napkin that's two axes: across, and up.

Pretend, just for the sketch, that we could name those two axes. Call one food. Call the other legal.

Notebook sketch of two arrows from the same origin. Pizza recipes points along food. Privacy policy points along legal.

"Pizza recipes" points along food. "Privacy policy" points along legal. They're almost at a right angle, which is another way of saying they don't have much in common. A person reading the two phrases would say the same thing.

A real embedding is this idea with hundreds of axes instead of two, and none of them come with names. You never get a chart like this. You only get how close the arrows sit.

Those arrows also have length. A longer page can make a longer arrow. Cosine similarity mostly shrugs at the length and asks about the angle, which is why a five-word query can still be compared with a 2,000-word page.

In plain terms

Think of it as the model pointing. Not summarizing the page. Not tagging the topic. Just pointing.

Why pages get turned into vectors

We used to need the same words. "How we price the starter plan" and "what the basic tier costs" can sit in a spreadsheet and never find each other. Embedding models are trained on a lot of text where those phrases hung around in the same neighborhoods, so they often point them the same way anyway.

Notebook sketch of two arrows pointing almost the same way along a price axis. Starter plan pricing and basic tier cost sit close together.

Same napkin. Different story. Both phrases are about what a plan costs, so both arrows lean along price. The angle between them is small. That's "close in vector space": not because they share a keyword, because they point the same way.

That's the swap. Embeddings took the comparison job from overlap metrics. You turn each page into a vector and ask how the arrows sit.

Semantic analysis is the next motion: group the nearby ones and see which rooms of the site are talking about the same work.

You'll hear a lot of people talk about "what a page means." We say it at Silkra too. It's sort of true. What meaning actually means to a model is the longer version of that caveat. Pages that point the same way often are about the same thing, which is why the shortcut is useful. Technically the model didn't understand anything. It ran a calculation, and we're the ones reading meaning into the result. The page got placed, not understood.

That placement still comes from training, plus whatever text you fed it. Titles alone make a jittery little arrow. Main content is usually the one you want. Dump the whole render, nav and footer included, and sibling pages start huddling because they share the chrome.

Same URL can produce three different arrows: title only, main content, or the whole page including the footer. If you don't know which text went in, the neighbor list is hard to read. We check that first.

What you can ask once pages are vectors

Once every URL has a direction, some of the tedious questions get cheaper.

  • Which pages point near each other even though they live in different folders?
  • Which page sits closest to a question a buyer would type?
  • Which passage is the nearest match, not which URL?

You could answer all of that with a crawl export and a long afternoon. Vectors get you to the interesting rows faster. They don't get you out of opening the pages.

A tight cluster might be a healthy topic group. It might also be three stubs saying the same thing in slightly different clothes. The arrow doesn't know. You do, once you look.

Same shape for retrieval. Embed the question, rank the passages by direction, and the winner is only the nearest available arrow. Nearest is not the same as good.

The inverse is useful too. Which pages sit far from everything else? Those meaning-space orphans often still look tidy in a folder tree. Vectors didn't invent that problem. They made it harder to miss.

How answer engines use these arrows

Answer engine optimization and generative engine optimization — AEO and GEO, if you need the acronyms — are mostly about this step. A model does not walk the sitemap. It embeds the question, embeds passages it might use, and pulls the ones pointing the closest way. Then it writes from what it pulled.

That is the same geometry as the napkin sketches. A buyer question about pricing is a short arrow. The passage that answers it is a longer one. If they point near each other, that passage is in the running. If the only close arrow is a changelog, the changelog is what shows up.

This is why vectors matter for the work. They are how you see topical similarity on a site: which pages would sit in the same neighborhood, which guide is next to the product page, which question has no nearby passage at all. We use them for that. It is a good use.

What they are not is a ranking lever you tune. You cannot stuff synonyms until the arrow looks more retrievable. You also cannot treat one embedding map as a replica of ChatGPT or Gemini. Those systems use their own models, their own way of cutting a page into passages, and a lot of steps after the nearest-arrow pass. A cluster view of a crawl is a picture of this text under this model. Useful. Not a crystal ball.

The practical version of AEO, if you strip the acronym off, is still writing a passage that answers the question, then checking whether that passage is the one that sits closest. Vectors help you do the check. They do not replace the writing. For a beginner walkthrough of that pipeline — and the parts we are guessing — see what happens when ChatGPT answers a question.

What the numbers leave out

This is the part decks usually skip.

Two pages can sit close because they share a brand voice, a product name, or a template. That is not the same as doing the same job for a searcher. Close is a hint. It isn't a topic label.

Nothing in the list tells you how many people want this. Keyword research still sees volume. Embeddings never did.

And the model didn't read the page. It placed the text. Useful, and a little dumb in ways you start to recognize: legal pages drift toward other legal pages, leftover nav copy pulls siblings together, product names and SKUs go a bit mushy.

You also can't optimize toward a vector by stuffing synonyms. The arrow moves when the words change. The work is still the words.

How this shows up in an audit

In the actual job it's quieter than the vocabulary.

Crawl the site. Pull the text you trust. Embed that, then look at who sits next to whom.

Two commercial pages stacked on each other? Open both and decide if they compete. A guide sitting next to a product page with no in-body link? That's a linking question that happened to show up in meaning-space. A pricing query that surfaces a changelog? Retrieval problem. Go inspect it.

None of those calls live in the numbers. The numbers beat scrolling a 4,000-row crawl table hoping the titles match.

Silkra's semantic map is one way of seeing the neighbors. We built it because staring at a spreadsheet of scores is miserable. The map still isn't the point. "Vector" in an SEO conversation almost always means: we turned the page into a direction so we could ask who it sits next to.

The part that is still fuzzy

We keep meeting SEOs who've been handed a cluster view and asked to treat it like a site architecture. It isn't one. Architecture is a decision someone made. A vector is a measurement of text.

What we haven't settled is the mix. How much of the audit is the measurement, and how much is still a person looking at the page. The measurement is getting cheaper. Swap the model and the neighborhoods move. If you're still in the loop, that's interesting. If you're not, it's a bad merge recommendation with extra confidence.

If "vector" has been floating around in a deck with no definition attached, this is the one we use: a list of numbers that points in a direction. Everything else — clusters, similarity scores, retrieval — is a way of comparing those directions.

Put crawl evidence to work

Download Silkra and turn audits, briefs, and fixes into one focused workflow.

Get started

Create your first workspace.

Crawl a site, then ask what needs attention.

Free to start. No credit card needed.