Customer Impact

SEO

Crawl budget explained: when it matters and when it doesn't

Copy for AI

Crawl budget is the number of pages a search engine like Google fetches on your site within a given period. The short answer to whether you should worry about it: for most B2B sites, no. Crawl budget only becomes a factor when your site is so large, so slow or so messy that Google no longer gets around to all your important pages. In this article you will learn what crawl budget actually is, where the threshold lies and when you are better off putting your attention elsewhere.

What is crawl budget exactly?

Before a page can rank, Google first has to fetch it (crawl it) and then include it in the index. Crawl budget is about that first step. It is not a fixed number Google assigns you, but the result of two things coming together.

The first is your crawl capacity: how many requests your server can handle without becoming slow or returning errors. If your site responds quickly, Google dares to fetch more pages per visit. If your server struggles, Google backs off so as not to overload your site.

The second is crawl demand: how much value Google places on your content to fetch it. Popular pages, pages that change often and pages with many links pointing to them get visited more frequently. A static page that never changes and that nothing links to sinks to the bottom of the queue.

Together, those two determine how much of your site Google actually looks at in practice. Crawl budget is part of technical SEO and belongs in the bigger picture of SEO, but it is a topic that in B2B often gets more attention than it deserves.

Who is crawl budget not a concern for?

Let us start honestly, because it will save you a lot of wasted time. If your site has a few dozen to a few hundred pages, crawl budget is almost never your problem. Google fetches sites of that size fully and effortlessly. It would be wasted energy to tinker with crawl budget while your content, internal links or loading speed are the real levers.

Most B2B service providers, product companies and local players fall into this category. A site with a homepage, a handful of service pages, an about page, a contact page and a knowledge base of a few hundred articles will never come near a crawl limit. If your pages are not in Google, the cause is almost always something else: a forgotten noindex tag, a page without a single internal link, or content that is simply too thin to be included.

In other words: if someone tells you that crawl budget is the reason your fifty-page site does not rank, you are entitled to be sceptical. That diagnosis is almost never correct.

From what threshold does it become a factor?

Crawl budget becomes relevant when the scale or the quality of your site starts to get in Google’s way. There is no exact number above which a switch flips, but a few signals help you gauge whether you are in the risk zone.

The first is sheer size. From roughly thousands to tens of thousands of URLs, the order and frequency in which Google fetches your pages start to matter. Think of large webshops, extensive knowledge bases, portals or sites that automatically generate many pages. From that point on, it can happen that new or changed pages are picked up more slowly because Google has to divide its attention.

The second signal has nothing to do with size, but with clutter. A smaller site can still work itself into trouble by producing huge amounts of low-value URLs. Common offenders:

  • Filter and sort URLs that generate endless combinations of parameters. How to tackle those specifically is covered in our article on faceted navigation and SEO.
  • Internal search result pages that are crawlable.
  • Session IDs or tracking parameters that show the same page under dozens of URLs.
  • Endless pagination or calendar pages that run years into the future.

In those cases, Google wastes its attention on pages that do not matter, so your real pages come up less often. The third signal is speed: a slow server or many server errors squeeze your crawl capacity, regardless of how big your site is.

EXAMPLE · WHERE YOUR CRAWLS GO Crawl attention that leaks away Filter & sort URLs 6000 URLs Internal search results 3000 URLs Session & tracking parameters 1800 URLs Real commercial pages 400 URLs What you actually want indexed Example figures for illustration
Junk URLs eat up the crawl attention your real pages need.

How do you know if your site is ready for it?

You do not have to guess. The crawl stats report in Google Search Console shows how Google has crawled your site over the past period: how many requests per day, how your server responded and what response times Google encountered. Two questions give you the most to hold on to.

The first: does Google reach your most important commercial pages and fetch them regularly? If so, there is little to worry about, regardless of how big the total number is. The second: do you see many server errors or slow response times? Those are the real alarm bells, because they slow Google down regardless of your size.

A second practical check is the ratio between the number of URLs your site generates and the number of pages you actually want to see indexed. If those diverge widely, for example because filters and parameters produce thousands of variants of a few hundred real pages, then you have a clutter problem that presents itself as a crawl budget problem. The solution then does not lie in more crawl capacity, but in cleaning up what Google gets to see.

What do you do about it when it does matter?

Suppose your site is large enough or produces enough clutter to make crawl budget a real factor. Then the approach comes down to a simple principle: let Google spend its attention on what matters, and not on what does not. A few concrete directions:

  • Clean up low-value URLs. Block crawlable filter combinations, internal search results and parameter variants that have no unique value, for example via your robots.txt or by not linking to them internally.
  • Use canonical tags consistently. When the same content exists under multiple URLs, use a canonical to point to the version that counts. This prevents Google from wasting time on duplicates.
  • Keep your sitemap clean. Put only the pages you actually want indexed in your XML sitemap, so you give Google a clear priority list.
  • Improve your server speed. Faster response times increase your crawl capacity, because Google dares to fetch more without overloading your site.
  • Strengthen your internal links. Important pages that get linked often are crawled more often. Deeply buried pages fade away.

If you cannot quickly implement those technical interventions via your CMS, then edge SEO offers a way to roll out redirects, canonicals and directives at the CDN layer without waiting for your dev team. And if you want to make changes to titles and metadata provable rather than based on gut feeling, look into SEO A/B testing.

Note that most of these measures make your site better anyway, regardless of crawl budget. That is exactly the right mindset: you solve an underlying quality problem, and better crawling is the consequence, not a goal in itself.

Frequently asked questions about crawl budget

Does my small B2B site have a crawl budget problem?

Almost certainly not. Sites with a few hundred pages or fewer are crawled fully and effortlessly by Google. If your pages are not in the index, the cause is almost always somewhere else, such as a noindex tag, missing internal links or content that is too thin.

From how many pages does crawl budget become important?

There is no hard limit, but as a rule of thumb it starts to play a role from roughly thousands to tens of thousands of URLs, or sooner when your site produces a lot of junk URLs and slow response times. Size and quality weigh together.

Does publishing more often increase my crawl budget?

Indirectly it can help, because an active site creates crawl demand. But you do not force a higher budget with it. A faster server, a clean URL structure and strong internal links usually have more effect than publishing frequency alone.

Is crawl budget the same as indexing?

No. Crawling is fetching a page, indexing is including it in the search results. Google can crawl a page and then still decide not to index it, for example because it adds too little value.

Spend your attention on the right lever

Crawl budget is a real mechanism, but for the vast majority of B2B sites it is not a problem that deserves your time. It only becomes a factor at significant scale, with many junk URLs or slow servers. The art is to first honestly establish whether it applies to you, rather than building a solution for a problem you do not have.

We are a small team that moves fast and honestly says what does and does not pay off. We steer on pipeline and revenue, not on vanity metrics, and we make sure you not only rank in Google but also get cited by AI search engines like ChatGPT and Perplexity. If you want to know where your site is really leaving gains on the table, our SEO specialist will map that out for you. Start with a look at the basics through our pillar on what SEO is.

Curious where your growth is getting stuck? Book your free intake

Free website scan

Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.

Where should we send your report?

We only use your details for your scan. No spam, unsubscribe anytime.