What happens when ChatGPT answers a question
A beginner sketch of how an LLM finds sources and writes an answer. What we can see from the outside, what to be skeptical of, and what we do not know.
Michael Davis/6 min read
When ChatGPT answers a question, it is easy to picture a person opening tabs. That is not what is happening. Something closer: a model that already knows a lot of text, plus, sometimes, a search over passages, plus a write-up of whatever it pulled.
We do not have the official pipeline. OpenAI does not publish the recipe. Google does not publish Gemini's. The systems change. What follows is the shape we can see from the outside. Treat it as a field sketch, not a spec.
It already knows a lot before it looks anything up
An LLM is trained on a huge pile of text, then frozen. That pile is not "the live web." It is a snapshot from some earlier window, plus whatever the lab added later. When you ask a question, a lot of the answer can come from that frozen pile. No crawl of a client site required.
That is why a model can talk fluently about a brand it has never fetched today, and why it can also be confidently wrong about a page that changed last week.
Training is not retrieval. Training is memory. It is also not a ranking of your URLs. If the only thing in the answer came from training, there is no "source page" in the SEO sense. There is a guess assembled from patterns.
Then, sometimes, it looks things up
Some answers stay inside the model. Some don't. When the system is allowed to look things up — a web search, a browse, a tool, a document someone pasted — a second step shows up.
The usual shape looks like this:
- Turn the question into a vector.
- Cut candidate pages into passages. A passage is a chunk of the page, not the whole URL.
- Turn those passages into vectors too.
- Pull the ones pointing the closest way.
- Write the answer from that mix of memory plus pulled text.

The question is a short arrow. The page has been cut into passages. Retrieval picks the nearest one, not the whole URL.
Search teams call the middle of that retrieval. Answer engines and "AI search" products are mostly this, with a lot of extra ranking and filtering we cannot see.
Sources can also arrive in blunter ways. The user pastes a URL. A browse tool fetches a page. A search API returns ten blue links and the model reads snippets. Those are different doors into the same idea: get some text in, then generate.
A citation is not a receipt. Sometimes the model used that page. Sometimes it used a snippet. Sometimes it named a URL that merely looks plausible. From the outside, those three cases can look identical.
The kind of content that tends to get used
This is the part people want as a ranking factor. It isn't one we can publish with a straight face. We can talk about patterns that show up often enough to be useful, and flag them as patterns.
Passages that answer the question. Retrieval is closer to "which paragraph is nearest the ask" than "which URL is the category page." A 2,000-word pillar that never states the answer can lose to a short section that does.
Text the system can actually extract. If the useful sentence only exists after JavaScript, or behind a login, or in an image, a lot of pipelines never see it. You cannot retrieve what you did not get.
Something distinctive enough to be the nearest arrow. Generic "what is SEO" copy sits in a crowded neighborhood. A specific explanation, a number you measured, a named process — those are easier to be close to, because fewer pages say them.
Not "long" as a virtue. Length can help if it adds a passage the question can match. It can also bury the answer under throat-clearing. Retrieval does not give points for word count.
None of that is "ChatGPT's algorithm." It is what you would expect if the job is "find nearby passages, then write." When we map a crawl by meaning, we are using the same idea to see topical neighbors on a site. That is a good use of vectors. It is not a replica of ChatGPT.
Five things this is not
These show up in decks a lot. They do not survive contact with the sketch.
It crawled the client site this morning.
Maybe a search tool fetched a page. Maybe it didn't. Training memory is not a live crawl.
A citation means that URL was the source of truth.
It means the system chose to show a link. That is a different job from "this is the document it read."
There is a ChatGPT ranking factor you can tune like PageRank.
There are many systems, they change, and the labs do not publish the weights. Anyone selling the official recipe is guessing with a price tag.
More schema / more synonyms will make the arrow 'more retrievable.'
Markup can help a machine find text. Synonyms can help a passage sit nearer a question. Neither is a lever with a published scale.
A GEO score is what ChatGPT retrieved.
A third-party score is one model's picture of some text. Useful as a sensor. Not a window into another company's stack.
Skepticism here is not "vectors don't matter." Vectors are how a lot of this lookup works. The skepticism is about treating one dashboard as the inside of ChatGPT.
What you can still do on a client site
Strip the acronyms off and the job is smaller than the conversation around it.
Write the passage that answers the question a buyer would ask. Put it in text a crawler can extract. Check whether that passage is the one sitting closest, or whether a changelog, a footer, or three older blog posts get there first.
Answer engine optimization and generative engine optimization — AEO and GEO — are names for that check. They are not a new ranking system you install.
If you want the geometry underneath, cosine similarity is how two arrows get compared. If you want the object being compared, start with the vector post. Neither one tells you what ChatGPT did on Tuesday. Both help you see whether a site has a passage ready for the questions that matter.
This is a sketch on purpose
We will be wrong about some of the internals. The products will move. A year from now a "source" might mean something stricter, or sloppier.
The useful part of writing this down is the gap it names. Training is not a crawl. A citation is not a receipt. A cluster view of a site is a picture of this text under this model. That picture is still worth making. It is how you find topical overlap, missing answers, and pages that would compete for the same question.
Just don't confuse the picture with the lab.
Put crawl evidence to work
Download Silkra and turn audits, briefs, and fixes into one focused workflow.
Get started›