SEO
Crawl budget optimization at scale (100,000+ URLs)
Copy for AI
On a site with more than 100,000 URLs, crawl budget is no longer an academic concept but a daily constraint. Google does not have the time or the inclination to fetch every page equally often, and if your structure leaks, Googlebot wastes its budget on filter pages and parameter variants while your new product pages stay out of the index for weeks. In this article you will read how to approach crawl budget optimization at scale: how to tame parameters, wall off crawl traps and use link sculpting to steer your crawl value toward the pages that generate revenue.
This is a part of SEO where many large sites quietly bleed without noticing. For the broader context you can first read our pillar on what SEO is; here we go deep for sites that are genuinely large.
When crawl budget actually starts to matter
Let’s start honestly. Google itself has repeatedly said that most websites do not need to worry about crawl budget. If you have a hundred-page B2B brochure site, Google crawls it without effort and tinkering with crawl budget is a waste of time. You win more there with strong content and good technical SEO than with crawl micromanagement.
Crawl budget only becomes a serious topic when you are into the tens or hundreds of thousands of URLs: large webshops with many variants, marketplaces, job sites, real estate portals or database-driven sites that generate pages from filters and searches. That is where the real problem arises: Google sets a crawl limit (how much your server can handle) and a crawl demand (how important and fresh Google considers your pages). Go above that and pages fall by the wayside.
You recognize the symptom like this: new or updated pages are indexed slowly or not at all, while in the meantime Google wastes thousands of requests on URLs that do not matter commercially. How many requests Google spends on what, you read off in the crawl stats report in Search Console. That report is your starting point for every optimization at scale.
The real leak: parameters and faceted navigation
On large sites the biggest crawl leak almost never sits in your actual content. It sits in the combinatorial explosion of URLs that your system generates without anyone realizing it.
Think of faceted navigation. A product category with filters for color, size, brand, price and sorting quickly produces thousands of unique URL combinations that all show virtually the same products. ?color=blue&size=l&sort=price and ?sort=price&size=l&color=blue are two different URLs with identical content to Google. Multiply that across hundreds of categories and you have millions of crawlable variants that swallow your real budget.
How do you tackle that?
- Decide which parameters are allowed to be indexable. A filter that yields a valuable, unique landing page (for example a brand within a category that people actively search for) may stay. Pure sorting and display parameters may not.
- Use a canonical to the clean category URL for filter combinations that have no search demand of their own. That way you consolidate signals instead of fragmenting them.
- Block worthless parameter combinations in robots.txt when they are crawled en masse without any indexing value. Note: a URL blocked via robots.txt can still end up in the index as a bare listing, so combine this thoughtfully with canonicals and internal links.
- Keep your internal links clean. If your navigation and breadcrumbs link only to canonical URLs, you prevent yourself from fueling the parameter sprawl.
The principle: don’t let Google figure out on its own which of your thousand variants is the real one. That costs crawl budget and creates confusion. Point it out.
Crawl traps: where Googlebot runs in circles
A crawl trap is a part of your site that in theory generates infinitely many URLs, so that Googlebot keeps pumping budget into it without ever finishing. They are notorious because they are stealthy: you only notice them once you analyze your crawl logs.
The classics on large sites:
- Infinite calendars. An events or booking calendar with a “next month” link that clicks endlessly through to the year 2147. Googlebot follows it dutifully, month after month.
- Filtered search results without a limit. Internal search URLs that make every query combination indexable, including the combinations nobody ever types.
- Session IDs and tracking parameters in URLs. Every visitor gets a unique URL, so every page exists in a thousand variants.
- Relative link errors that stack paths infinitely, such as
/category/category/category/.... - Pagination that never stops or keeps generating pages beyond the last real results.
The approach starts with measuring. Analyze your server logs or the crawl stats report and look at where Googlebot sends its requests. Often you then see that a hefty share of your crawl budget goes to a handful of URL patterns that have zero commercial value. You shut those patterns down: block them in robots.txt, put noindex on pages that must remain reachable but should not enter the index, and remove or nofollow the links that feed the trap. For calendars you simply limit how far forward and back one can click.
Honestly, this is detective work that pays off. Every wasted crawl you win back is a crawl that can go to a page that does bring in customers.
Link sculpting: steer crawl value toward your money pages
Once you have plugged the leaks, the second part is about giving direction. Google crawls and reindexes pages more often the better they are linked internally and the closer they sit to the homepage. That is the lever of link sculpting: arranging your internal structure so that crawl and link value flows toward your most important commercial pages.
In practice that means:
- Make your most important pages shallowly reachable. A page that sits five clicks deep is rarely crawled. Ensure your money pages sit within two or three clicks of the homepage via category hubs and internal links.
- Build thematic hubs that link to the underlying pages and vice versa, so that crawl value keeps circulating within the cluster instead of leaking away to filter pages. Our guide on internal links goes deeper into that structure.
- Prune links to low-value URLs. Every link your navigation places to a sorting page tells Google that page is important. Deliberately reduce those links in favor of pages that matter.
- Keep your XML sitemap clean and current. Put only canonical, indexable URLs in it. A sitemap full of redirects, 404s and parameter variants undermines your own priority signal.
Link sculpting is no longer a trick with nofollow attributes like ten years ago. It is architecture: determining which pages produce pipeline and tuning your entire internal link structure to them. That is exactly the kind of work an experienced SEO specialist does when a large site has plenty of pages but gets too few of the right ones into the index.
Optimize for the right outcome, not for the total
The biggest pitfall with crawl budget is steering on the wrong number. A high number of crawls per day feels good, but it says nothing if those crawls go to the wrong pages. At Customer Impact we treat SEO as the acquisition layer of your growth engine, optimized for pipeline and not for vanity dashboard numbers. That applies here too.
The question that matters is: does Google reach and reindex your commercially important pages on time? If you launch a new product line or service and it stands fresh in the index within days, then your crawl budget is working for you. If it is still out of the index after three weeks while Googlebot scours thousands of filter pages, then you have a budget problem that directly costs revenue.
And don’t forget that the same discipline also helps you with AI search engines. A clean, well-structured site that guides crawlers efficiently through your most important content is not only indexed better by Google but also more easily picked up and cited by ChatGPT, Google AI and Perplexity. Crawl efficiency is a foundation, not a side issue.
Get started at your scale
Optimizing crawl budget on a large site is not a one-off cleanup but an ongoing discipline: taming parameters, closing crawl traps and steering your internal structure so that your money pages get priority. The difference between a site that indexes fast and one that lags months behind rarely lies in more content and almost always in a cleaner, better-directed architecture.
Is your site running into indexing problems, or do you suspect Google is wasting its budget on the wrong URLs? Get in touch and we will look together at where your crawl budget leaks and how we send it back to the pages that produce pipeline.
Free website scan
Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.
We only use your details for your scan. No spam, unsubscribe anytime.