SEO
How to find orphan pages: the valuable pages nobody links to internally
Copy for AI
An orphan page is a page on your site that no other page links to. The short answer: because crawlers travel the web by following links, they barely find such a page, it gains almost no internal authority and it stays invisible in the search results. The annoying part is that it is often your most valuable pages that end up orphaned: an old service page, a strong article, a product page that dropped out of the menu during a redesign. In this article you will learn how to track them down systematically using crawl, log file and sitemap data, and how to decide what should happen to them.
What exactly is an orphan page?
Search engines discover pages largely by following links. From your homepage, a crawler hops to your menu, to category pages, to articles and service pages. Every page hanging in that network gets found. An orphan page sits outside that network: not a single internal link points to it, so the crawler has no path to reach it.
That is something different from a poorly linked page. A page with one single link buried deep in your site is weakly positioned, but still reachable. A true orphan page has zero internal links. It does exist, you can open it through the direct URL, but within your internal link structure it simply is not there.
Why does this happen so often? The most common causes are a website migration where links were not carried over, products or pages that disappeared from a menu or filter, imported content that was never linked, and standalone landing pages for campaigns that live outside the main structure. The bigger and older your site, the more orphans quietly appear. Programmatic SEO, where you generate pages at scale, also produces orphan pages quickly if you do not build the internal links in right away.
Why a regular crawl does not find them
Here is the pitfall. Tools like Screaming Frog crawl your site exactly the way Google does: they start on the homepage and follow links. But precisely because an orphan page has no internal links, such a crawl cannot reach it by definition. The page you want to find is invisible to the method you are using.
The solution is to compare the crawl against sources that know your pages through a channel other than internal links. An orphan page still appears in your XML sitemap, in your analytics if it gets visits, and in your server logs if a bot or a visitor ever opened it. By laying those lists next to your crawl, the pages that do exist but are missing from the crawl stand out.
Find orphan pages with sitemap data
The fastest first step is your XML sitemap. Your sitemap is in principle your own complete inventory of the pages you want indexed. Many SEO crawlers let you load the sitemap URLs alongside the regular crawl. The tool then crawls your site and reads your sitemap, and sets both lists against each other.
What you are looking for are the URLs that appear in the sitemap but do not show up in the crawl. Those are pages you considered important enough to put in the sitemap, but that nothing on your site links to. That is your first and most reliable list of suspected orphan pages. Keep in mind that your sitemap itself has to be current and complete, otherwise you miss exactly the pages that are not in it.
Find orphan pages with analytics and Search Console
Your analytics package and Google Search Console know pages that were visited or shown at some point, regardless of your internal links. Export the list of URLs that received visits or impressions over the past months, and compare it with your crawl.
A page that pulls in traffic or impressions but does not sit in your crawl is doubly interesting: it is already performing despite being orphaned. That is often a page you can easily grow further by pointing a handful of internal links at it.
With this source, watch out for one detail. Analytics only knows pages that received visits in the recent period. A valuable page that has not attracted traffic for a while, precisely because it is orphaned, can therefore stay under the radar. That is why you always combine analytics with your sitemap and your logs: only when you bring the three sources together do you get a complete picture of what lives outside your structure. How you make sure pages end up in the index at all is something you can read in our article on getting your website indexed.
Find orphan pages with log file analysis
The most thorough source is your server logs. A log file records every request to your server: which URL was requested, by which bot or visitor, and when. That shows you the raw reality of how Googlebot crawls your site in practice, not how you think it does.
To track down orphan pages you do two things. First, you pull all unique URLs that were requested from the logs and compare them with your crawl: pages that sit in the logs but not in the crawl are orphan candidates. Second, you see which pages Googlebot rarely or never visits, which is tied to how your crawl budget is distributed. Orphaned pages generally get little bot traffic, because there are no links sending the bot their way. How to read that bot traffic is something we explain in our article on the crawl stats report in Search Console.
Log file analysis takes more work than a sitemap export and requires access to your server logs, which not every host exposes by default. For a small site it is often overkill and a sitemap plus analytics will do. For large sites with thousands of pages, where crawl budget genuinely matters, it is the only way to see with certainty which pages Google ignores in practice. So scale your effort to the size and the stakes of your site.
Not every orphan belongs back in your structure
Before you start linking: not every orphan page is a mistake. Some pages are orphaned on purpose, and that is exactly right. Think of thank-you pages after a form, separate landing pages for a paid campaign, or pages you explicitly want to keep out of the organic index. Those are not pages you want to suddenly start linking from your menu or your content.
Assess every page you find on intent. Broadly, they fall into three categories. Valuable pages that were orphaned by accident, you pull back into your structure with relevant internal links. Outdated or thin pages without value are better merged, rewritten or retired via a redirect. Deliberately orphaned pages you leave alone, possibly with a noindex so the signal matches. That judgement call is exactly where a good SEO specialist makes the difference: not blindly linking everything, but choosing what contributes to pipeline.
From finding to linking
For the pages that do need to come back, the fix is often surprisingly simple. Look for existing pages that are thematically related and add a natural, contextual link there. A few relevant links from strong, existing content still give the page crawl paths and internal authority. A logical spot in your website architecture makes sure the gain lasts instead of being a one-off.
The order of your optimisations matters. First find them via sitemap, analytics and logs. Then judge them on intent. Only then link, merge or retire. Anyone who starts laying links straight away without checking the value pollutes their structure all over again.
At Customer Impact we do not treat orphan pages as a standalone technical clean-up, but as part of one growth engine: SEO is the acquisition layer that not only gets you ranking, but also gets you cited in AI search engines like ChatGPT and Google AI. The goal is never a neat list of resolved orphans, but pages that genuinely contribute to enquiries. You can read more about that approach on the pillar page about what SEO is.
Want to know which valuable pages are hanging invisibly in the margins of your site? Get in touch and we will map out your orphaned pages and their potential.
Free website scan
Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.
We only use your details for your scan. No spam, unsubscribe anytime.