What is internal linking? Meaning, and what changed for AI search

The short answer
Internal linking is the practice of linking your own pages to each other. The definition has not changed. The job it does has.
For twenty years it did two things. It helped crawlers find pages, and it passed authority between them.
It now does a third. An agent assembling an answer can follow your links to gather context. A controlled 2026 study measured what that is worth.
Key takeaways
- Internal linking is still a discovery and authority tool. That part of the old advice holds.
- It is now also a retrieval path. Agentic systems traverse links to collect what they need before answering.
- A controlled study measured the gain. Enhanced entity pages with rich interlinking improved accuracy by 29.6 percent in standard retrieval.
- Depth still matters. A page nobody can reach in a few clicks gets crawled rarely, and read as unimportant.
- Descriptive anchors beat generic ones. “Read more” tells a machine nothing about what sits on the other side.
Where this comes from. We run organic for more than 250 enterprises at Pepper and track over 10 million prompts across every major engine. We audit link structures most weeks, and the same faults recur. That work shapes the view here, along with the client reviews we sit in and the talks we have at the events we run.
What is internal linking?
Internal linking means placing a link on one page of your site that points to another page on the same site.
That is the whole definition. What matters is what those links do.
The three jobs
- Discovery. A crawler finds pages by following links. A page with no links pointing at it is hard to find and easy to miss. Crawler access gates this before links do.
- Authority. Links pass ranking value between pages. A page linked from many strong pages tends to rank better than an orphan.
- Retrieval context. This is the new one. An agent answering a question can follow your links to gather related facts before it writes.
The first two are classic SEO. The third is why this subject is worth revisiting in 2026.

See where you show up. Pepper’s GEO platform tracks Brand Visibility, Domain Prompt Presence and Share of Voice across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews. Your team can log in, connect Search Console and GA4, manage the prompt set and run your own agents in the Agent Atlas. A growth team works the same account alongside you. Book a growth audit or see where you show up.
What changed for AI search
A paper submitted on 11 March 2026 tested this directly, and it is the clearest evidence we have found.
Andrea Volpini, Elie Raad, Beatrice Gamba and David Riccitelli ran seven conditions. Three document formats, crossed with two retrieval modes, plus one extra variant. The formats were plain HTML, HTML with JSON-LD, and an enhanced entity page carrying rich interlinking. The retrieval modes were standard RAG, and agentic RAG with multi-hop link traversal.
They tested four domains: editorial, legal, travel and ecommerce. The stack was Vertex AI Vector Search and Google’s Agent Development Kit.
The result
The enhanced entity page format improved accuracy by 29.6 percent in standard retrieval, and 29.8 percent in the full agentic pipeline. The strongest variant reached 4.85 out of 5 for accuracy and 4.55 for completeness.
So link structure is not only a crawl concern now. It changes how well a machine can answer a question using your site.

Read the limits before you act on this
Three of them matter, and the authors are open about two.
- This is a controlled experiment rather than a measurement of production engines. It shows what helps a retrieval pipeline. It does not prove ChatGPT or Google Search behave the same way.
- The incremental gain from the richest variant was not statistically significant over the base enhanced format, which the authors state plainly.
- The abstract does not give a corpus size. We would want that before calling the exact percentage precise. So we treat the direction as the finding.
One further note on a related claim. The enhanced format in that study included llms.txt-style instructions. That is not evidence that publishing llms.txt helps in production AI search. The server-log evidence shows 97 percent of those files go unread. A pipeline you design, and a search engine you do not, are different things.
The faults we find most often
Four, and they account for most of the damage we see in audits.

Orphan pages
A page with no internal links pointing at it. Crawlers find it late or never, and it inherits no authority.
These usually appear after a migration, or when pages ship from a template that nothing links to. A content audit will surface them alongside the pages worth keeping.
Pages buried too deep
Click depth is how many links a visitor follows from the homepage to reach a page. Pages sitting four or more clicks down get crawled less often. They also read as less important.
Most sites can fix this with better hub pages rather than a rebuild.
Generic anchor text
“Click here” and “read more” describe nothing. The anchor is one of the few places you tell a machine what the target page is about, so wasting it costs you.
Write the anchor as the thing the reader will get. How engines read page text is covered in our analysis of 1.4 million prompts.
Links that go only one way
A cluster where every page links up to the hub, and the hub links back to none, loses most of the benefit. Relationships should be visible from both ends.
How to do internal linking well
Six rules. None of them needs a tool.
- Link from your strongest pages. A link from a page that already ranks carries more than one from a page nobody reads.
- Keep important pages within three clicks. Use hub pages to pull deep content up rather than restructuring everything.
- Write anchors that describe the target. The anchor should make sense read on its own, out of the sentence.
- Link both ways inside a cluster. Hub to spoke, and spoke back to hub. Topical authority covers how to choose the subjects worth clustering.
- Link where it helps the reader. If the link would not help a person mid-sentence, it will not help a machine either.
- Vary your anchors. The same anchor text on every link to one page reads as mechanical, and it gives engines less to work with.
How many links per page
There is no magic number, and anyone quoting one is guessing. We aim for links that earn their place. In a long article that usually means three to eight contextual links, plus whatever navigation already provides.
Our methodology: how we weight internal linking work
| Factor | Weight | Why it carries this much |
|---|---|---|
| Click depth of money pages | 30% | Depth predicts crawl frequency, and a page crawled rarely is retrieved rarely |
| Descriptive anchor text | 25% | One of the few direct statements you make about a target page |
| Cluster completeness, both directions | 25% | Retrieval agents traverse links, and one-way clusters break the path |
| Orphan page elimination | 20% | A page nothing links to inherits nothing and is found late |
**What we could not verify and left out.** The widely repeated claim that AI engines preferentially cite passages reachable within two internal link hops of the homepage. It appears across several SEO blogs with no study behind it, and we could not find a primary source. Also excluded: any specific optimal number of internal links per page, for the same reason.

Internal linking approaches at a glance
| Approach | What it delivers | Typical 2026 cost | Who it suits | Where it falls short |
|---|---|---|---|---|
| Manual contextual links | Relevance, and the best anchors | Minutes per page, ongoing | Most sites | Does not scale past a few hundred pages |
| Hub and spoke clusters | Topical structure a machine can follow | A planning day, then ongoing | Content-led sites | Needs maintenance as the cluster grows |
| Breadcrumbs | Cheap depth reduction, and clear hierarchy | A developer afternoon | Every site, especially deep ones | Structural only, no topical signal |
| Related-posts automation | Volume, with no human time | Plugin or build cost | Large publishers | Relevance is often poor, anchors generic |
| Full link audit | Orphans, depth and broken paths named | A crawl plus a day of reading | Sites over a few hundred pages | Produces a list, not the fixes |
| Platform plus a growth team | Structure, plus the work it points at | Custom | Teams without internal capacity | Still gated by your publishing pace |
—
How to evaluate your internal linking
A crawl answers most of this in an afternoon.
The scorecard
Score your site out of 100.
| Factor | Weight | How to score it honestly |
|---|---|---|
| No orphan pages that matter | 25% | You have crawled the site and checked, rather than assuming |
| Money pages within three clicks | 25% | Measured from the homepage, with the actual numbers written down |
| Anchors describe their targets | 20% | You read twenty at random and none of them said “read more” |
| Clusters link both ways | 20% | Hub to spoke and back, checked on your three main clusters |
| Links serve the reader first | 10% | Someone can defend each link as useful mid-sentence |
The live test, over 90 days
Crawl the site and write down three numbers: orphan count, average click depth, and the depth of your ten most valuable pages.
Then fix the worst of them and leave everything else alone, so you can read the result.
Fix 30 prompts in your buyers’ words and run them monthly across Google, ChatGPT and Perplexity for 90 days. Real ones, such as “what is internal linking”, “best data observability tool for a small team”, “how much does sales enablement software cost” and “how do I fix orphan pages”.
Record whether you are named and whether a page of yours is cited. Three readings is the shortest honest window. One is a snapshot, two could be noise, three shows direction. Keep the set fixed, because changing it resets the comparison.
Watch the deep pages you lifted. If they start getting cited and nothing else changed, the structure was the constraint rather than the writing.
One caution on reading the result. Structural fixes move slowly, because a crawler has to return before anything changes. Give it the full ninety days before judging.
Weak approach versus strong approach
The weaker approach: install a related-posts plugin, let it generate links, then count the total as progress. The anchors will be generic and the relevance poor, so the links say nothing about what they point at.
The stronger approach: crawl once and fix orphans. Lift your ten most valuable pages to within three clicks, then write anchors a person could read on their own.
One adds links. The other adds meaning, and only the second changes what a machine can do with your site.
Red flags
- A target number of internal links per page. No study supports one, and the people quoting numbers do not cite any.
- Automated links with generic anchors. They add nothing a machine can read.
- “Two hops from the homepage” quoted as a rule. We could not find a primary source for it.
- Link audits that end at a spreadsheet. The list is not the work.
- Links added for machines, not readers. If it does not help a person, it is noise.
- Guaranteed ranking gains from internal links. Nobody controls generated output, and Google warns against guaranteed rankings.
Five questions worth asking
- “How many orphan pages do we have?” If nobody has crawled, nobody knows.
- “How deep are our ten most valuable pages?” Measured, not guessed.
- “Would these anchors make sense read alone?” The fastest quality check there is.
- “Do our clusters link both ways?” One-way clusters break the traversal path.
- “Which links would you remove?” A good answer means somebody has actually read them.
Reduced to one principle: link so a person is helped and a machine can follow.
One closing note that costs us something. Internal linking is largely free, it needs no platform, and most sites can fix the worst of it in a day with a crawler and an afternoon of reading, ours included in the general point that this is easier to sell as a project than it is to justify as one.
What does internal linking cost?
| What you are buying | Typical 2026 cost | What it covers |
|---|---|---|
| A crawl of your own site | Free to low | Orphans, click depth, broken links |
| Reading the results | An afternoon | Which pages are buried and which anchors are useless |
| Fixing anchors | Minutes per page | The cheapest meaningful change here |
| Hub pages to reduce depth | A day or two | Lifting deep pages without a rebuild |
| Breadcrumbs | A developer afternoon | Structural depth, site-wide |
| Ongoing linking discipline | Minutes per new page | Keeping it from degrading again |
| Platform plus a growth team | Custom | Pepper: tracking, plus the people and agents doing the work |
The first two rows are free or near-free, and they tell you whether the rest is needed.
What nobody should promise you
A specific number of internal links per page. No published study supports one.
Ranking gains from link volume alone. Links carry meaning through anchors and placement, not through count.
That internal links alone will get you cited. They help a machine reach and understand a page. The page still has to be worth quoting.
Guaranteed results. Nobody controls generated output, and Google warns against providers guaranteeing rankings.
When internal linking is not your problem
The answer that costs us the sale. If your site has fewer than about fifty pages and everything sits two clicks from the homepage, you do not need internal linking work and we would tell you not to buy it.
Small sites are shallow by default. The structural problems here appear with scale, usually past a few hundred pages or after a migration. Spend the time on whether the pages are worth citing instead. That is a writing question rather than a structural one.
Frequently asked questions
What is internal linking?
Linking one page of your site to another page on the same site. It helps crawlers discover pages, passes ranking value between them, and now also gives retrieval agents a path to follow when assembling an answer.
Why does internal linking matter for AI search?
Because agentic retrieval traverses links. A 2026 controlled study found an enhanced entity page format with rich interlinking improved retrieval accuracy by 29.6 percent over plain HTML in standard retrieval.
How many internal links should a page have?
There is no supported number, and anyone quoting one is guessing. Aim for links that earn their place. In a long article that usually means three to eight contextual links alongside your normal navigation.
What is click depth?
The number of links a visitor follows from the homepage to reach a page. Pages sitting four or more clicks deep get crawled less often and tend to be read as less important by search systems.
What is an orphan page?
A page with no internal links pointing at it. Crawlers find it late or not at all, and it inherits no authority from the rest of the site. Orphans usually appear after migrations, or from templates nothing links to.
Does anchor text still matter?
Yes, and more than link count. The anchor is one of the few places you state directly what the target page covers, so “read more” wastes the opportunity entirely.
Should I use automated internal linking tools?
With care. They add volume quickly, but the anchors are usually generic and the relevance poor. On most sites, a short manual pass over the twenty most valuable pages beats thousands of automated links.
How is internal linking different from backlinks?
Internal links connect your own pages and you control them completely. Backlinks come from other sites and you do not. Both help discovery, but only internal links are yours to fix this afternoon.
Where to go next
Crawl your own site and write down two numbers: how many orphan pages you have, and how deep your ten most valuable pages sit.
Those two numbers tell you whether this is worth a day of your time. On most sites past a few hundred pages, they are worse than anyone expects.
Learn AI Search · Run the technical audit · Book a growth audit
Sources and further reading
Primary studies
- Andrea Volpini, Elie Raad, Beatrice Gamba and David Riccitelli. “Structured Linked Data as a Memory Layer for Agent-Orchestrated Retrieval”, submitted 11 March 2026. Seven conditions: three document representations, being plain HTML, HTML with JSON-LD, and an enhanced agentic-optimised entity page, crossed with two retrieval modes, being standard RAG and agentic RAG with multi-hop link traversal, plus an Enhanced+ condition adding rich navigational affordances and entity interlinking. Four domains: editorial, legal, travel and ecommerce, using Vertex AI Vector Search and Google’s Agent Development Kit. The enhanced entity page improved accuracy by 29.6 percent in standard RAG and 29.8 percent in the agentic pipeline, with Enhanced+ reaching 4.85 out of 5 for accuracy and 4.55 for completeness. The authors state the incremental gain of Enhanced+ over the base enhanced format is not statistically significant, and the abstract does not give a corpus size. This is a controlled retrieval experiment rather than a measurement of production search engines. Link
- Ahrefs. Study of 1.4 million ChatGPT prompts. 88.46 percent of citations came from pages retrieved through the search channel, which is why discovery and crawlability still gate everything downstream. Link
Further reading
- Google Search Central. AI search optimisation guidance, 15 May 2026. Confirms AI Overviews and AI Mode run on core Search ranking and quality systems, and warns against providers guaranteeing rankings.
- Pepper. What is llms.txt, and should you use it for why a controlled pipeline and a production engine are different things, and what AI engines actually cite for the page-level half.
- Deliberately excluded: the claim that AI engines preferentially cite passages within two internal link hops of the homepage, which appears across several SEO blogs with no study behind it. Also excluded, any specific optimal number of internal links per page.
Latest Blogs
The research pulls in two directions. Fan-out rewards covering many adjacent questions; absorption research finds longer pages get used. The resolution is that the passage competes, not the page.
Seven documented changes, not seven predictions. Google switched two rich results off, shipped two new AI reports, and added a channel to Analytics. Here is what actually moved and what it means.
Content chunking is the process of splitting a page into smaller pieces so a machine can retrieve one of them. Here is the part most articles get wrong: chunking is something engines do to your page, not something you do to it. You do not choose the chunk size.
Get your hands on the latest news!
Similar Posts

SEO
15 mins read
Primary and secondary keywords in 2026: what they were always for, and what changes now

Content Creation
11 mins read
What is a SERP? A complete guide to search engine results pages in 2026

SEO
15 mins read