What is semantic HTML for SEO and AEO?
Semantic HTML names the parts of a page: main, article, nav. What those tags do for people, extractors, SEO, and answer engines — and what they do not do.
Silkra team/8 min read
Semantic HTML is markup that names the parts of a page. <main> is the primary content. <nav> is navigation. <article> is a self-contained piece. A <div> is a box with no name.
Search extractors, screen readers, and answer-engine crawlers all have to guess where the article is. Named landmarks make that guess smaller. They do not rank the page by themselves, and they do not make ChatGPT cite you.
What semantic HTML is
HTML has two jobs. It can describe how a page looks, or it can describe what a region *is*. Semantic elements do the second job.
The useful set for this conversation is small:
`<main>`
— the primary content of the document. One per page.
`<article>`
— a self-contained composition: a post, a product write-up, a docs page.
`<nav>`
— a block of navigation links.
`<header>` / `<footer>`
— introductory or closing matter for the page or a section.
`<section>`
— a themed grouping, usually with a heading.
`<h1>`–`<h6>`
— the outline, in order.
MDN's note on semantics is the dry version: elements that carry meaning about their content, not just a hook for CSS.
A page built from anonymous <div>s can look identical. A person using a mouse may never notice. A machine trying to pull "the article" out of "the chrome" notices every time.
ARIA landmarks (role="main", role="navigation") can name the same regions when you cannot change the tags. They are a backup, not a reason to skip the elements.
In plain terms
Semantic HTML is labeling the rooms. Div soup is an open-plan office with no signs.
What it changes for a crawler
A crawler, a reader-mode parser, and a chunker all start from the same problem: which bytes are the page, and which bytes are the template?
When <main> or <article> is present, a lot of extractors treat that region as the candidate. When it is missing, they fall back to heuristics: biggest text blob, Readability-style scoring, "strip the nav if we can find it." Those heuristics fail in familiar ways. The cookie banner wins. The related-posts rail wins. The footer of legal links wins. The extracted "content" is chrome.
We see this in audits as a thin page that is not thin. The live URL has a 1,200-word explanation. The extract has 80 words of navigation. What we check before trusting a crawl is mostly this class of lie.
Semantic HTML does not guarantee a good extract. A <main> that also wraps the whole site chrome is a nametag on the wrong door. A page that loads the article after JavaScript, with an empty <main> in the first HTML, still hands the crawler a shell if it does not wait.
The markup is a hint. Rendering and extraction still have to do the work.

One side has a name on the article. The other is boxes all the way down.
Does it matter for SEO?
For classic search, the honest answer is quieter than the 2014 blog posts. Google has said for years that it can understand pages without a perfect outline. Valid HTML helps machines; it is rarely the lever that moves a competitive query by itself.
Where it still shows up:
Accessibility is certain. Screen readers use landmarks and heading order to skip around. Skipped levels and missing <main> are real problems for people. They are also cheap to fix.
Extraction is practical. Title, meta description, and the text Google chooses to snippet all depend on what the system decided was the document. If your unique paragraph never made it into that document, you are optimizing a different page than the one you published.
[Headings](/blog/do-headings-still-matter) are part of the same story. An H2 is a suggested tear line for people and for some chunkers. It is not a keyword slot.
We have not run the clean test this topic wants: same prose, two markup structures, compare what Google and a handful of extractors pull. Until that exists, we will not pretend we have a ranking coefficient for <article>. A named main region is easier to extract and easier to defend in an audit. That is enough to write the tags.
Does it matter for AEO?
Answer engine optimization and generative engine optimization — AEO and GEO — care about whether a passage can be found and used. That starts before any vector comparison.
If OpenAI's crawlers or a browse tool fetch the URL and the useful sentence lives only in a client-rendered widget, or only in an image, or only outside <main> in a template the extractor discarded, the passage never enters the pile.
Named regions help the "get the text" step. They do not help the "pick this URL over a competitor" step. Nobody has published a ranking factor called "uses <article>."
The failure mode we keep seeing is the opposite of missing tags: the tags are there, and they wrap the wrong thing. <main> around the entire app shell. <article> on every card in a grid. An H1 in the nav. The names become noise, and the extractor is back to guessing.
So the AEO version of the standard tags is still boring. Put the answer in text. Put that text in a region that is actually the main content. Give it a heading that names the question. Then check the extract, not the source file.
There is a second story in the SEO conversation, and it is a different job.
Extra labels beyond the standard tags
Some SEOs now talk about labeling more than <main> and <nav>. aria-label on an icon button. A name that tells two navs apart. A form field that says what it wants, not just "Source." Custom data- attributes, HTML comments, even little glosses on a <section> meant to "steer" a model that is fetching or clicking through the page.
Two jobs get mixed together here. One is retrieval: cut the article out, embed the passages, maybe cite the URL. The other is an agent using the page: click Add to cart, fill the form, open the right tab. Extra names matter much more for the second job than for the first.
OpenAI says so, for the agent. Their publishers and developers FAQ is blunt: ChatGPT Atlas uses ARIA tags — "the same labels and roles that support screen readers" — to interpret page structure and interactive elements. They ask for descriptive roles, labels, and states on buttons, menus, and forms so the agent can tell what each control does. Google's agent-UX note describes the same map: agents look at screenshots, raw HTML, or the accessibility tree, and that tree is roles, names, and states.
That is vendor guidance about using a site. It is not a claim that extra adjectives in aria-label change who OAI-SearchBot cites.
What the studies actually measured
There is research that models can read markup, not just the visible sentence. Gur, Nachum, and colleagues at Google (Understanding HTML with Large Language Models, Findings of EMNLP 2023) treated HTML understanding as its own task: classify what an element is, describe a control, navigate a page. Fine-tuned models transferred better than systems trained only on the HTML task. Structure is in the input. That paper does not test whether aria-label="enterprise pricing calculator" on a <div> makes a citation more likely.
The computer-use work is easier to over-read. A11y-CUA (Berkeley and Michigan, CHI 2026) ran an agent on 60 desktop and web tasks. Default success was about 78%. Keyboard-only dropped to about 42%, magnifier to about 28%. Agents fail when they lose the easy visual path. That is not a study of extra labels lifting AEO citations.
The caution from accessibility still applies. The first rule of ARIA is to use a native element when one exists. A real <button> already has a role. A real <label for="…"> already names the field. WebAIM's Million surveys have long shown that pages reaching for ARIA tend to carry more errors, not fewer, because the attributes get bolted onto markup that already lied. A wrong name is worse for an agent than a missing one. The agent treats the tree as ground truth.
So the extra-labeling narrative is interesting, and half of it is real. Label the controls an agent has to use, honestly, the same way you would for a screen reader. Distinguish "Main navigation" from "Footer navigation" when both exist. Do not build a parallel taxonomy of marketing phrases on every <section> in the hope of steering a search crawl. We have not seen a measurement that this moves citations. We have seen it make the accessibility tree noisier.
What we look at on a client page
A short pass, before anyone argues about models:
- Is there a
<main>orrole="main"? - Does that region contain the unique paragraph, or the chrome?
- Do the headings outline the article, in order, without skips?
- After a rendered crawl, does the extract match the live page?
- Do interactive controls have an honest name — visible text, a
<label>, oraria-labelwhere native HTML cannot say it?
Silkra flags pages with no semantic container because we had to clean the whole <body> to get text. That is an extraction warning, not a ranking verdict. The fix is usually a landmark around the article you already wrote. Names on buttons are a separate pass, for agents and for people.
The part that is still fuzzy
Answer-engine decks now ask whether <article> is an AEO factor, or whether a richer set of labels will steer the model. The studies we can point to are about extraction quality, HTML understanding, and agents moving through a UI. They are not about extra adjectives in aria-label changing who gets cited.
The measurement we still want is the controlled pair: same words, two trees, compare the extract and whether an agent can finish a task. Until that exists, write the names for the rooms. Name the controls that do something. Check the extract. Leave the "steer the LLM" layer for someone who has run the test.
Put crawl evidence to work
Download Silkra and turn audits, briefs, and fixes into one focused workflow.
Get started›