Internal Link Graph Design: The Numbers That Decide Your Site Architecture

Internal link graph design is the practice of deciding, page by page, which URLs on your site link to which, with what anchor text, and at what depth from the entry point. Most teams treat it as a cleanup task. It is a capital allocation decision: every internal link you place moves crawl attention and ranking authority from one page to another, and the pages you starve stay invisible whatever their quality. This pillar gives you the metrics that prove or disprove an internal link graph design, the studio's house position on where the field is wrong, three engagements where the playbook was load-bearing, and the order in which to read the rest of this hub. You will finish able to defend the call in a board meeting, with numbers.
Key takeaways:
- Google's link documentation states that every page you care about should have a link from at least one other page on your site. Pages failing that test are orphans, and Ahrefs traces most of them to migrations, redesigns and discontinued products.
- Zyppy's analysis of 23 million internal links across 1,800 websites, updated 23 February 2026, found URLs with 0 to 4 internal links averaged 2 Google clicks, URLs with 40 to 44 internal links averaged roughly four times that, and the curve reverses past 45 to 50 links.
- Pages carrying at least one exact match internal anchor had at least five times the search traffic of pages without one, in that same dataset.
- PageRank has always been a division problem: the 1998 Brin and Page paper sets the damping factor at 0.85 and splits a page's score across its outbound links.
- Google's crawl budget guidance only engages above 1 million pages changing weekly, or 10,000 pages changing daily. Below those counts, link graph work buys ranking and discovery, not crawl capacity.
- Internal link graph design is measured on the destination side: inbound contextual links, distinct anchors and depth per URL, never the count of links you happen to place on a source page.
What internal link graph design actually decides
Internal link graph design is the discipline of choosing which pages on a site link to which others, with what anchor text, and at what distance from each entry point. The output is a directed graph: pages are nodes, links are edges, and search crawlers traverse it to discover URLs, to read anchor text as topical context, and to compute how much authority each node inherits. Google is unambiguous about the discovery half of that job. Its link best practices documentation states that a link is only crawlable when it is an HTML anchor element carrying an href attribute, and that every page you care about should have a link from at least one other page on your site. The authority half traces to the founding paper. Sergey Brin and Lawrence Page defined PageRank in 1998 with a damping factor of 0.85 and a score divided across every outbound link a page carries, which makes each new link a decision to dilute the ones already there.
Internal link graph design therefore operates on three levers at once: which edges exist, what those edges say, and how far each node sits from a real entry point. Move any one of the three and the other two change. This hub answers a single question: how do you decide the shape of a site's graph, and which number tells you the decision was right? Every topical, focal and special entry underneath it inherits that question and narrows it to one surface, one topology or one failure mode. The hub sits inside the Growth, SEO and content engineering family, alongside programmatic publishing, schema, analytics and attribution. Internal link graph design is the layer that decides which of those pages a crawler ever reaches.
Three systems consume the graph, and they consume it differently. The crawler follows edges to find URLs it does not know about. The indexer reads anchor text and treats it as a description of the destination written by a third party. The ranking layer computes flow across the whole graph, so a change to one navigation template propagates to every page that template renders. Treating those three as one problem is where most architecture work goes wrong, because a fix that helps discovery, such as a 400-link HTML sitemap, actively hurts flow. Internal link graph design keeps the three separate and prices each proposed fix against the specific metric it moves.
The studio's house position on internal link graph design
A link graph is a budget you spend. Every page carries a finite score, that score divides across its outbound links, and adding a link to a page reduces what every other link on that page transmits. Teams that treat internal linking as a coverage exercise, adding related-posts modules and tag clouds until no page is unlinked, spend that budget uniformly across pages of wildly unequal commercial value. The field's standard advice, restated across the Moz blog and most technical SEO guidance, correctly says that internal links pass equity and establish hierarchy. It stops short of the harder claim. Internal link graph design starts from the division, not from the coverage.
Here is where we disagree with the field. Link count targets are the wrong unit. The recommendation to place 5 to 10 internal links per article, repeated in nearly every checklist, prescribes a quantity on the source page when the number that predicts traffic sits on the destination page. Zyppy measured inbound links per URL rather than outbound links per article, and found the relationship monotonic up to a ceiling and then negative. An article with 10 outbound links contributes nothing if all 10 point at pages already saturated with inlinks.
Second disagreement: crawl budget is invoked far below the threshold where it exists. Google's crawl budget guidance scopes itself to large sites of 1 million or more unique pages with content changing about once a week, and to medium or larger sites of 10,000 or more unique pages with very rapidly changing content daily. Google adds that these numbers are a rough estimate to help you classify your site rather than exact thresholds. A 400-page marketing site has no crawl budget problem. It has a distribution problem, and the two demand opposite fixes: crawl budget work removes URLs, distribution work moves authority toward the URLs that convert. Internal link graph design on a site of that size is entirely a distribution exercise, and every hour spent on robots.txt there is an hour not spent on the graph.
Third, navigation is not the graph. Sitewide navigation and footers link every page from every page, which flattens the graph into something close to uniform and destroys the signal that tells a crawler which page owns a topic. Contextual body links carry the differentiation. Zyppy's controlled tests on selective link priority found that when a page carried two text links to the same destination, Google indexed the first text anchor and ignored the second, so a boilerplate nav link placed above your carefully written body link can consume the anchor slot you wanted. The studio's rule is blunt: a page's topical position is set by the body links pointing at it, and nav links are plumbing.
The consequence for planning is that internal link graph design happens before you write, not after you publish. Deciding retroactively which of 300 articles deserves inbound support is a retrofit, and retrofits are how sites end up with programmatic SEO surfaces generating thousands of URLs that the graph never reaches.
Five numbers that prove or disprove your architecture
An architecture claim you cannot measure is an opinion. These five metrics turn the claim into something a board can audit, and every entry in this hub reports against at least one of them. Internal link graph design is defensible only when each of the five carries a current number and a stated target.
| Metric | How to measure it | Healthy range | Failure signal |
|---|---|---|---|
| Click depth to revenue pages | Crawl from the homepage, record the shortest path to each URL | 3 clicks or fewer | Money pages sitting at depth 5 or deeper |
| Inlinks per published URL | Count distinct internal followed inlinks per destination | 10 to 45 | A median of 1, or a long tail above 50 |
| Distinct anchors per destination | Group inbound anchor text per URL | 3 or more variants, at least one exact match | One boilerplate anchor repeated everywhere |
| Orphan ratio | Sitemap URLs minus crawl-reachable URLs, over total | Under 2% | The crawler finds more orphans than in-structure pages |
| Graph coverage | Share of published URLs receiving a contextual body link | Above 95% | Navigation links carrying the whole graph |
Click depth measured from every entry point
Depth from the homepage is the default report in every crawler, and it flatters sites whose traffic never lands on the homepage. Measure depth from each page that actually receives external links and organic entries, then take the minimum per URL. A page three clicks from a homepage nobody enters through, and six clicks from the guide that earns your backlinks, is a depth-six page in practice. Recomputing depth this way moves part of the archive into a worse band than the homepage-only report showed, and that delta is the number worth reporting to whoever signs off on the architecture.
Inlink distribution and the reversal past 45 links
The Zyppy dataset covering 23 million internal links across 1,800 websites and roughly 520,000 URLs is the strongest public evidence on this metric. URLs with 0 to 4 internal links averaged 2 clicks from Google Search. URLs with 40 to 44 internal links averaged around four times that. Past 45 to 50 inbound internal links, average traffic declines as the count rises. Report the distribution rather than the mean: a site averaging 18 inlinks per URL can still have 60% of its pages sitting at 1, which is the shape internal link graph design exists to correct.
Anchor variety per destination
Cyrus Shepard, the author of the Zyppy SEO study, states the finding plainly: "Pages with at least one exact match anchor had at least five times more traffic than pages without." He adds that "anchor text variety is highly correlated with higher search traffic." The same study found pages linked with naked URL anchors received roughly 50% more traffic than pages without them, and that over 6% of all links carried no anchor text at all with no measurable traffic difference. Google's own guidance sets the qualitative bar: read the anchor text out of context and check whether it is specific enough to make sense by itself.
Orphan ratio and the crawl it wastes
Orphan pages are pages with no internal links from elsewhere on the site, which means crawlers reach them only through sitemaps or external links. Botify's published worked example shows how far this runs on large sites: Googlebot crawled roughly 800,000 orphan URLs against 300,000 pages in the site structure, so more than 70% of what Google explored sat outside the architecture, while just 5% of organic visits came from those orphans and 95% came from the 30% of pages that were in-structure.
Internal link graph design coverage as a publishing gate
Coverage is the share of published URLs that receive at least one contextual body link from another page. Making it a publishing gate rather than an audit metric is the single highest-leverage change most teams can make. No page ships until a specific existing page links to it in prose, with an anchor written for the destination. Internal link graph design that lives only in an annual audit regresses inside two publishing quarters, because every new page ships unlinked and the orphan set refills faster than the audit clears it.
Hub and spoke, mesh, or flat: picking a topology you can defend
Three topologies cover almost every real site, and internal link graph design picks between them from the content inventory rather than from taste.
| Topology | Choose it when | Depth profile | Main risk | Decision rule |
|---|---|---|---|---|
| Hub and spoke | Content divides cleanly into 5 to 15 topics with a clear owner page per topic | 2 to 3 clicks | Spokes talk only to the hub, so lateral relevance is lost | Default for editorial archives under 2,000 URLs |
| Mesh | Entities relate many-to-many, such as products by attribute or locations by service | 2 to 4 clicks | Uncontrolled edge growth pushes pages past the 45-link reversal point | Only with a generated link layer and a cap per page |
| Flat | Fewer than roughly 150 URLs, all of comparable commercial value | 1 to 2 clicks | No hierarchy signal, so no page owns a topic | Small sites and single-product surfaces |
Hub and spoke is the studio default, with one amendment the standard model omits: spokes link to each other, not only up to the hub. A pure star gives the hub every inbound edge and leaves each spoke with one, which lands the entire spoke inventory at the bottom of the inlink distribution described above. Adding two to four lateral links per spoke moves the median without touching the 45-link ceiling.
Mesh topologies fail in one specific way. Generated related-content modules compute similarity and emit links without a cap, so popular attribute pages accumulate hundreds of inbound edges while the source pages emit fifty or more outbound links each, diluting every one of them below usefulness. Cap outbound generated links at a fixed number, ranked by commercial value rather than by similarity score, and the topology holds.
Flat architectures are defensible below roughly 150 URLs and indefensible above it. The tell is a navigation menu that has grown to 30 or more items because there is nowhere else to put the links.
The topology choice interacts with how answer engines read a site, which is covered under generative engine optimisation. Answer engines extract passages rather than pages, so a topology that makes each page a self-contained answer on one sub-topic outperforms one that spreads a topic across five loosely joined URLs. Internal link graph design for answer engines is an ownership problem before it is a flow problem.

Three engagements where internal link graph design was load-bearing
The insights archive you are reading. The strategy behind laboetie.io carries 1,894 planned entries mapped to 228 SERP clusters across ten families, of which 73 paths are live at the time of writing. An archive of that size built without a link contract becomes an orphan farm within two quarters, so the internal link graph design was fixed before the first article shipped: every entry declares its family, its hub and its parent, the hub pillar links every published spoke, and no entry publishes without a body link from an existing page. Sibling entries link laterally rather than routing every path through the hub, which is the amendment described in the topology section above. The published llms.txt file mirrors the same shape, listing eight sections and 18 canonical links, with the insights archive described as an encyclopedia of digital sovereignty organised into ten thematic categories. The engine running it is Cortex, the studio's SEO and GEO growth engine, whitelist only and capped at three brands per industry.
A French savings and retirement publisher. france-epargne.fr is a finance archive, which means every page is a Your Money Your Life page and every page competes against banks and insurers carrying far heavier domain signals. The load-bearing decision was to refuse a flat blog and commit to one hub per product family, with the hub owning the head term and spokes owning the qualifiers. In finance the cost of getting this wrong is asymmetric: a product page sitting at depth five behind a tag archive competes for a term worth thousands of euros per conversion and loses to a competitor whose equivalent page sits at depth two. The gate that made it hold was editorial rather than technical, since every new guide had to name the product page it supported before it entered the queue.
An insurance comparison surface. assurecompare.fr and the Lynkflow platform behind it face the mesh problem in its purest form, because insurance products relate to each other by carrier, by guarantee, by professional status and by price band. Uncapped generated linking here produces exactly the failure the topology table describes, with a handful of popular attribute combinations absorbing the whole graph. Internal link graph design here is a capping problem: a hard limit on generated outbound links per page, the ranking of which links survive the cap set by commercial value rather than by similarity score, and a contextual body link required from an editorial page before any comparison surface enters the graph at all.
Across all three, the studio's sovereignty thesis applies to the graph itself. Clients keep ownership of the architecture and the link layer, which means no part of the graph may depend on a vendor plugin the client cannot export.
The eight-point link graph audit you can run this week
Run these in order. Each produces a number, and the numbers are the argument for or against your current internal link graph design.
- Full crawl against the sitemap. Crawl the site from the homepage, export every reachable URL, and diff against every URL in the XML sitemaps. The difference is your orphan set. Anything in the sitemap the crawler never reached has no internal link, which is the definition Ahrefs uses.
- Inlink distribution histogram. Bucket every published URL by its count of distinct internal followed inlinks, in bands of five. Read where the mass sits. A left-heavy histogram with a median of 1 or 2 is the most common finding on sites over 500 pages.
- Depth from real entry points. Recompute click depth from the top 20 pages by external links and by organic entries, not from the homepage alone, and record the minimum per URL.
- Anchor text inventory per destination. For each commercially important URL, list every inbound anchor. Flag any destination whose anchors are all identical boilerplate, and any destination with no exact match anchor.
- Nav versus body ratio. Split inbound links per destination into template links and body links. Any URL whose inbound count is 100% template is not positioned on any topic, whatever the raw count says.
- Outbound link count per template. Count outbound internal links on each page template. Templates emitting more than 45 to 50 links sit past the point where the Zyppy curve reverses, and every additional link is dilution.
- Redirect and broken edge sweep. Internal links pointing at redirects or 404s waste both the crawl and the flow. Report them as a percentage of total internal edges; above 2% is a maintenance failure.
- Coverage gate compliance. For every URL published in the last quarter, verify that a contextual body link existed at publication. This is the forward-looking control, and it is what turns internal link graph design into a standing constraint rather than a project you repeat next year.
What is changing in this hub this year
The graph now has a second audience, and it reads differently. Answer engines assemble responses from passages pulled across sources, so the unit of citation is a section rather than a page. Otterly.AI's analysis of more than 1 million AI citations across ChatGPT, Perplexity and Google AI Overviews from January to February 2026 found brand domains taking 47.5% of citations against 52.5% for community platforms such as Reddit and Quora, with the brand share varying sharply by platform: 59.8% on Google AI Overviews, 44.7% on ChatGPT and 28.9% on Perplexity. The same analysis reports that 73% of sites carry technical barriers blocking AI crawler access, which makes crawler reachability a prerequisite before any graph work pays.
Two practical consequences follow. Internal link graph design now has to tell an answer engine which page owns a topic, because that engine will cite one passage and attribute it to one URL. Ambiguous ownership, where four pages each cover 30% of a topic, splits the citation across four weak candidates and none of them wins. Second, agent crawlers do not execute JavaScript, so a link rendered client-side does not exist for them, and a site whose navigation is a client-only component has no graph at all from their perspective.
The file layer is the other change. A published llms.txt gives agents a curated map of the site that sits beside the link graph rather than replacing it, and the studio covers that surface in the agent-facing web and llms.txt hub. Teams that want the file without the engagement can generate one directly with the studio's llms.txt Generator, which is self-serve, paid once, and needs no account.
What has not changed is the arithmetic. Damping at 0.85 and division across outbound links still governs how authority moves, and Google's own description of discovery still puts link following first among the methods it uses to find pages. Internal link graph design in 2026 adds an audience; it does not replace the mechanism.
Which entry in this hub to read first, by starting condition
Pick by the condition you are actually in, because internal link graph design decisions branch hard on the starting point.
- You inherited a site and do not know its shape. Start with the audit above, run points 1 to 3, and come back with an orphan ratio and an inlink histogram. Every other decision depends on those two numbers.
- You are choosing a topology before building. Read the topology table above, then size your content inventory. Under 150 URLs, stay flat. Between 150 and 2,000 with clean topic boundaries, build hub and spoke with lateral spoke links. Above that, or with many-to-many entities, plan a capped mesh, because internal link graph design at that scale needs a generated layer with a hard limit rather than editorial discipline alone.
- You are generating pages programmatically. The graph is the constraint, not the generator. The generation side belongs to the programmatic publishing hub in the same family; the cap on outbound generated links belongs here.
- You are optimising for answer engines. Pair this hub with llms.txt generation and platform setup, because citation follows passage ownership and passage ownership follows the graph.
- You want the whole picture. The full insights archive indexes every family and lists the sibling hubs that touch this one, including schema markup, analytics and attribution.
FAQ: internal link graph design
How many internal links should a page have?
Count inbound, not outbound. Zyppy's 23-million-link study found URLs with 0 to 4 internal links averaged 2 Google clicks while URLs with 40 to 44 averaged around four times that, with the curve reversing past 45 to 50 inbound links. On the outbound side, keep templates below roughly 45 internal links so each one still transmits meaningfully. Internal link graph design sets a destination-side floor of about 10 inbound contextual links for any page you expect to rank.
Does internal link graph design still matter for AI search?
Yes, for two reasons. Answer engines cite passages and attribute them to a single URL, so the graph decides which of your pages owns a topic and therefore which one gets cited. Otterly.AI measured brand domains at 47.5% of more than 1 million AI citations between January and February 2026, with wide variation by platform. Reachability comes first: the same analysis found 73% of sites carry technical barriers blocking AI crawler access.
What counts as an orphan page?
Ahrefs defines orphan pages as pages search engines may have difficulty discovering because they have no internal links from elsewhere on the website. They usually appear after migrations, navigation changes, redesigns, discontinued products, and test or development pages left in the sitemap. Detect them by diffing a full crawl against your XML sitemaps. Anything present in the sitemap and absent from the crawl is an orphan.
Should I worry about crawl budget?
Only above Google's stated scope: large sites of 1 million or more unique pages with content changing about once a week, or medium and larger sites of 10,000 or more unique pages with very rapidly changing content daily. Google calls these a rough estimate rather than exact thresholds. Below that, your problem is distribution of authority and discovery of new URLs, and the fixes differ: crawl work removes URLs, distribution work redirects internal links toward pages that convert.
Do navigation links count the same as body links?
They do not carry the same weight in practice. Sitewide navigation links every page from every page, which flattens differentiation and tells a crawler nothing about which page owns a topic. Zyppy's selective link priority tests found that when one page carried two text links to the same destination, Google indexed the first text anchor and ignored the second, so a nav link can consume the anchor slot your body link was written for.
What anchor text should internal links use?
Descriptive text written for the destination, with at least one exact match anchor among the inbound set. Google's guidance is to read the anchor out of context and check whether it is specific enough to make sense by itself. Cyrus Shepard's finding is the commercial argument for variety in internal link graph design: pages with at least one exact match anchor had at least five times more traffic than pages without.
How La Boétie works on internal link graphs
The studio treats architecture as an engineering deliverable with a specification, not as an ongoing retainer. Three parts, in the order they run.
Graph diagnostic. A full crawl, an inlink distribution histogram, depth recomputed from real entry points, and an anchor inventory per commercially important URL. The output is the five-metric table from this pillar, filled in with your numbers, plus the orphan set as a list of URLs. The eight-point audit above is the same procedure the team runs internally before proposing any change to an internal link graph design.
Architecture specification and build. A topology decision you can defend, a link contract every future page must satisfy, and the templates and generated link layers that enforce it. Because clients keep ownership of what gets built, the link layer ships as code in your repository, with no dependence on a vendor plugin you cannot export. The sovereignty thesis the studio takes from Étienne de La Boétie in 1548 applies to your architecture as much as to your stack.
Publishing gate and operation. Coverage becomes a gate rather than a report: no page ships without a contextual body link from an existing page, and the archive is instrumented so the gate is measurable each quarter. The same discipline runs this archive, where 1,894 planned entries across 228 SERP clusters share one link contract and a team of five to six engineers maintains it.
If you are carrying a site whose numbers you do not know, book a studio intro call. Bring the site and the question you need answered; the team comes back with the diagnostic, and with an opinion about what to build rather than an estimate for what you asked for.
Conclusion
An internal link graph is a finite budget divided across the pages you choose to support, and the division has been arithmetic since Brin and Page set damping at 0.85 in 1998. The field's default advice, more links and more coverage and more related-posts modules, spends that budget uniformly across pages of unequal value and calls the result a strategy. The five metrics in this pillar replace that with something auditable: click depth from real entry points, inlink distribution against the 45-link reversal, anchor variety with at least one exact match, orphan ratio under 2%, and coverage enforced as a publishing gate. Sound internal link graph design is what those five numbers look like when they are all inside their bands at once.
The decision rule is short enough to carry into a meeting. Measure the destination side rather than the source side. Choose the topology from the content inventory. Cap generated links by commercial value. Gate publication on a contextual body link. Applied in that order, internal link graph design stops being a maintenance chore and becomes the mechanism that decides which of your pages a search engine, and now an answer engine, is able to reach at all.
Sources
- 23 Million Internal Links, SEO Case Study : Zyppy SEO, Cyrus Shepard, updated 2026
- How Google's Selective Link Priority Impacts SEO : Zyppy SEO, Cyrus Shepard, 2025
- Link best practices for Google : Google Search Central, 2026
- Large site owner's guide to managing your crawl budget : Google Search Central, 2026
- The Anatomy of a Large-Scale Hypertextual Web Search Engine : Sergey Brin and Lawrence Page, Stanford InfoLab, 1998
- Orphan Pages and SEO : Botify, 2023
- How to Find and Fix Orphan Pages : Ahrefs, Kayle Larkin, 2022
- The AI Citation Economy : Otterly.AI, 2026
- Moz blog, internal linking and site architecture : Moz, 2026
Also read:
Questions
How many internal links should a page have?
Count inbound, not outbound. Zyppy's 23-million-link study found URLs with 0 to 4 internal links averaged 2 Google clicks while URLs with 40 to 44 averaged around four times that, with the curve reversing past 45 to 50 inbound links. On the outbound side, keep templates below roughly 45 internal links so each one still transmits meaningfully. Internal link graph design sets a destination-side floor of about 10 inbound contextual links for any page you expect to rank.
Does internal link graph design still matter for AI search?
Yes, for two reasons. Answer engines cite passages and attribute them to a single URL, so the graph decides which of your pages owns a topic and therefore which one gets cited. Otterly.AI measured brand domains at 47.5% of more than 1 million AI citations between January and February 2026, with wide variation by platform. Reachability comes first: the same analysis found 73% of sites carry technical barriers blocking AI crawler access.
What counts as an orphan page?
Ahrefs defines orphan pages as pages search engines may have difficulty discovering because they have no internal links from elsewhere on the website. They usually appear after migrations, navigation changes, redesigns, discontinued products, and test or development pages left in the sitemap. Detect them by diffing a full crawl against your XML sitemaps. Anything present in the sitemap and absent from the crawl is an orphan.
Should I worry about crawl budget?
Only above Google's stated scope: large sites of 1 million or more unique pages with content changing about once a week, or medium and larger sites of 10,000 or more unique pages with very rapidly changing content daily. Google calls these a rough estimate rather than exact thresholds. Below that, your problem is distribution of authority and discovery of new URLs, and the fixes differ: crawl work removes URLs, distribution work redirects internal links toward pages that convert.
Do navigation links count the same as body links?
They do not carry the same weight in practice. Sitewide navigation links every page from every page, which flattens differentiation and tells a crawler nothing about which page owns a topic. Zyppy's selective link priority tests found that when one page carried two text links to the same destination, Google indexed the first text anchor and ignored the second, so a nav link can consume the anchor slot your body link was written for.
What anchor text should internal links use?
Descriptive text written for the destination, with at least one exact match anchor among the inbound set. Google's guidance is to read the anchor out of context and check whether it is specific enough to make sense by itself. Cyrus Shepard's finding is the commercial argument for variety in internal link graph design: pages with at least one exact match anchor had at least five times more traffic than pages without.