What is chunking, and why do LLMs do it?

A chunk is a passage cut from a page so a model can retrieve it. Why pages get split, and what a mid-sentence cut does to the nearest answer.

Michael Davis7 min read

Summarize

A model does not retrieve a URL. It retrieves a chunk: a passage cut from the page, small enough to embed and compare with a question. Something has to decide where that cut goes. That decision is chunking.

When the cut is clean, a question can find the paragraph that answers it. When the cut lands mid-thought, the nearest passage is a fragment. The page can still be "there." The answer isn't.

Models retrieve passages, not pages

What happens when ChatGPT answers a question is mostly this shape: turn the question into a vector, turn candidate passages into vectors, pull the nearest ones, then write. The unit in the middle is not the page.

That is a constraint, not a preference. Embedding models have a context window. So do the steps that rank passages. A 2,000-word guide will not go in as one object. It gets sliced first.

So the SEO object and the retrieval object are different. You publish a URL. The system stores pieces. A buyer question matches a piece. If the piece that would have answered it was split across two slices, neither slice is a great neighbor for the question. The guide can still rank in a classic sense. The passage the model needed is gone.

This is also why "we have a page for that" is a weaker check than it used to be. You can have the URL. You can even have the sentence. If the sentence is not sitting inside a passage that still makes sense on its own, retrieval may never see it as the answer.

Answer engine optimization and generative engine optimization — AEO and GEO, if you need the acronyms — spend a lot of airtime on models and dashboards. A lot of the work is earlier and smaller: did anyone leave a complete passage?

What a chunk is

A chunk is a slice of a page that a retrieval system treats as one unit. Not the URL. Not the title tag. A passage.

How big that slice is depends on the pipeline. Some systems cut on headings: an H2 plus the text under it. Some use a fixed window of tokens, often a few hundred, sometimes with overlap so the end of one slice is repeated at the start of the next. Some mix both. The labs do not publish the recipe, so treat any "ChatGPT chunks at N tokens" slide as a guess.

The job of the slice is simple. It has to be short enough to embed. It has to be long enough to mean something. Those two pressures pull in opposite directions.

A 12-word chunk can be a heading and a leftover. A 1,200-word chunk can bury the sentence that would have matched. Overlap papered over some of the mid-sentence cuts. It also means the same words get stored twice.

You do not get to pick the other company's knife. You do get to look at a page and ask whether a reasonable knife would leave anything usable.

In plain terms

Chunking is cutting a page into the pieces a model can hold. Retrieval then picks a piece, not the whole URL.

We cut a real page on purpose

We took the published cosine similarity note — 1,411 words of body text — and ran a dumb splitter: 900 characters, no overlap, no respect for headings or sentences. That is not Silkra's production cut, and it is not ChatGPT's. It is the kind of window a lot of RAG tutorials still start with. Ten chunks came out.

The interesting failure is in the middle of the page, not the intro. The note's load-bearing sentence is the one about 0.85 looking huge and still being ordinary. The 900-character window ended one chunk at "routinely reached 0." The next chunk opened on "85. That is not 85% similar."

The number that does the teaching got split in half.

Earlier cuts were messier in a quieter way. Chunk 1 died inside the -1 / 0 / 1 list, on the letter *m* of "means." Chunk 2 started "eans they sit at right angles." A model can still embed that. What it cannot do is recover the list you meant to publish.

Notebook sketch of a page on the left and the same page cut into passages on the right. A steel slash lands in the middle of a word.

The charcoal cuts are the boring ones. The steel cut is the one that walks through a word.

If we had cut on H2s instead, that page is six sections plus a short lede. The definition section is 308 words. The "why it showed up" section is 247. Those are complete thoughts. They are also larger than a 900-character window, so a heading-aware split and a character window do not produce the same objects. Same URL. Different passages.

A bad cut is a retrieval problem

Embed "reached 0." and you get an arrow. It just is not an arrow about cosine similarity. It is an arrow about a fragment.

Cosine similarity then compares that fragment with a question. The question "what does 0.85 mean" wants the sentence that says 0.85 can look huge and still be ordinary. That sentence no longer exists as one object. Half of it lives with the previous thought. Half of it opens a new chunk with no subject.

Retrieval does not glue those back together and then decide. It ranks the pieces it has. A weak but intact paragraph on another URL can beat a broken one on the "right" page.

This is the part that is easy to miss in an audit. You open the live page. The explanation is there. You ask a model a question the page answers. It cites something else, or it cites the page and still mangles the point. The temptation is to blame the model. Sometimes the model is looking at a slice that never contained the point.

Overlap helps. A 100-character overlap would have carried "0.85" into both sides of that cut. Heading-aware cuts would have kept the whole section. Neither is free, and neither is what every system does.

The practical read: if a passage only works as an answer when you can see the sentence before it, it is a fragile chunk.

What you can still do when you write

You cannot install ChatGPT's splitter. You can write pages that survive a clumsy one.

Put a complete answer under a heading that names the question. If a system cuts on H2s, that section becomes a chunk that still makes sense. If it cuts on characters, you at least gave the window a self-contained run of sentences to land in.

Do not hide the useful sentence at the seam. A definition that starts at the end of one section and finishes at the start of the next is one bad window from disappearing.

Length is not a strategy. Extra words can add a passage the question can match. They can also push the answer into the next slice. We have watched both.

Extractable text still matters. A chunker can only cut what a crawler got. If the explanation lives in an image, a tab, or a render the extractor missed, there is no chunk to argue about.

None of that is a ranking lever. It is damage control for a step you do not control.

When we crawl a site in Silkra, we chunk the extracted text so we can search passages instead of titles. That is one picture of one splitter. Useful for seeing whether a question has a nearby passage at all. Not a replica of another company's cut.

The part that is still fuzzy

We do not know how each answer engine cuts a page on Tuesday. Window size, overlap, heading rules, whether the title gets prepended to every slice — those knobs move, and they are not public.

What we can see is the shape. Pages get cut because models retrieve passages. A clean passage can be the nearest arrow. A split one is two worse arrows. The cosine note we cut is a friendly example: we wrote it, we know where the point lives, and a 900-character window still walked through the number.

The question we keep coming back to is how much of "this page didn't get used" is a content problem and how much is a cut problem. Those look the same from the SERP. They do not look the same once you read the slices.

We want to run the same two cutters — heading-aware and a fixed window — across a handful of client pages next. One friendly blog post is enough to see the shape. It is not enough to know how often the load-bearing sentence is the one that gets split.

Put crawl evidence to work

Download Silkra and turn audits, briefs, and fixes into one focused workflow.

Get started

Create your first workspace.

Crawl a site, then ask what needs attention.

Free to start. No credit card needed.