SEO
Faceted navigation SEO: the filters eating your crawl budget
Copy for AI
Faceted navigation is the filter bar you see on virtually every webshop: filter by brand, size, colour, price, availability. For the visitor, that is convenient. For search engines, it is one of the most stubborn technical SEO problems in existence. Every filter combination generates its own URL, and those URLs multiply at breakneck speed. In this article you will read why faceted navigation SEO eats your crawl budget, and which choices to make with noindex, canonical and robots.txt to take back control.
If you want the foundation underneath all of this first, take a moment to read up on what SEO actually involves and how it fits into a broader growth strategy. This article builds on that with a specific e-commerce problem. The same filter explosion happens on job boards, incidentally, where filters for role, region and sector produce thousands of URLs; our approach there is covered in SEO for recruitment and staffing agencies.
Why filters cause a URL explosion
Say you have a webshop with 300 products. Manageable enough. But your product overview has five filter groups: brand (10 options), size (6), colour (8), price range (4) and availability (2). A visitor can click those filters in every conceivable combination, and each combination often gets its own URL with parameters, for example ?brand=x&size=42&colour=black.
The number of combinations is not a sum, it is a multiplication. Add parameter order, sort options and pagination on top, and you are easily looking at tens of thousands of unique URLs for a catalogue of a few hundred products. Google sees all of them as separate pages it can discover, crawl and potentially index.
That causes three concrete problems:
- Crawl budget waste. Googlebot has a finite number of pages per site that it crawls within a given period. If it burns that on thousands of filter combinations, your important product and category pages stay unvisited for longer, or sit unchanged in the index.
- Duplicate and thin content. Many filter combinations show virtually the same, or nearly empty, product lists. Google then has to decide which version is the real one, and that decision does not always fall in your favour.
- Diluted signals. Internal links and any external links spread across dozens of variants of the same list, instead of concentrating on the page you actually want to rank.
The three choices per filter type
There is no single button that fixes this in one go. The approach is a deliberate decision per filter type: do you let it be indexed, do you consolidate it, or do you block it entirely? Three instruments, three different roles.
1. Index: filters with real search demand
Some filter combinations match how people genuinely search. Think of a brand combined with a product category (“Nike running shoes”) or a size that is in high demand. Pages like that have commercial value and are welcome in your index.
To that you add two conditions. The page needs a static, clean URL rather than a parameter string, and it has to show enough unique content: its own title, a descriptive text, a well-stocked product list. Treat such filter pages as fully fledged landing pages, not as accidental filter results. How to build URLs like that is covered in our explainer on the SEO-friendly URL.
2. Canonicalise: variants of the same list
For filters that do produce a logical page but have no search demand of their own, you use the canonical tag. A sort order, a price filter or a colour filter then points back via rel="canonical" to the clean main category.
That tells Google: these variants exist, but the main page is the version you should rank. The canonical consolidates the signals onto one page. Note, though: a canonical is strong advice, not an order. Google may ignore it when the content diverges too much. And, important for crawl budget: Google still has to crawl the page to read the canonical. So you solve duplication, but you save no crawl time.
3. Block: noise nobody searches for
Sort direction, session IDs, tracking parameters and endless worthless filter combinations do not belong in a search engine. Those you block preferably before Google crawls them, via robots.txt. A rule like Disallow: /*?sort= keeps whole classes of parameter URLs out of the crawl.
This is exactly where the trade-off sits that many webshops misjudge. A URL blocked via robots.txt is not crawled, so you save crawl budget. But Google does not read the page either, and therefore sees no noindex or canonical on it. Block a page that is still linked elsewhere, and a bare URL without a description can still end up in the index. Robots.txt and noindex are not synonyms: one governs crawling, the other governs indexing.
Robots.txt, noindex and canonical: who does what
This is the part where most mistakes originate, so let us put the roles sharply side by side:
- robots.txt (Disallow) governs whether Google may crawl a URL. It saves crawl budget, but consolidates no signals and guarantees no exclusion from the index.
- meta robots noindex governs whether a crawled page may enter the index. It keeps the page out of the search results, but the page has to be crawlable for the tag to be read.
- rel=“canonical” governs which version counts as the real one. It consolidates duplicates, but saves no crawl time and is advice, not a guarantee.
The mistake you want to avoid: blocking a page in robots.txt and putting a noindex on it at the same time. Google can then never read the noindex, because it is not allowed to crawl the page. The result is the opposite of what you wanted. If you really want a page out of the index, leave it crawlable with a noindex, and only block it in robots.txt once it has disappeared from the index.
In practice, then, you use all three in layers. The bulk of the noise you block pre-emptively in robots.txt. The meaningful but duplicate variants you leave crawlable with a canonical. And the pages with real search demand you turn into clean, indexable landing pages.
How to check whether it works
You do not guess whether it is working, you measure it. In Google Search Console, the crawl stats show how much of your crawl budget goes to parameter URLs. If that gets out of hand, you will see it there. We wrote a separate explainer on how to read and interpret crawl stats in Search Console.
The Page indexing report then shows which URLs are excluded and why: “Crawled, currently not indexed”, “Duplicate without user-selected canonical”, “Blocked by robots.txt”. Those labels are your dashboard. If the number of duplicate and excluded parameter URLs rises, your filter strategy is not in order yet. If you work with a large catalogue in particular, this is one of the first things we look at when a webshop outsources its e-commerce SEO.
Why this reaches further than crawling
Setting up faceted navigation properly is not a standalone technical chore. It determines whether your crawl budget goes to your money pages or evaporates into filter noise, and whether your catalogue comes across as an ordered whole or as a sprawl of half-duplicates. That directly affects your organic visibility, and increasingly also whether AI search engines pick up your product pages as a reliable source instead of some random filter variant.
At Customer Impact we do not treat SEO as a set of isolated technical tricks, but as the acquisition layer of an orchestrated growth engine. Filter strategy, indexing and internal links: we build them so they produce pipeline, not just tidy graphs. Want an experienced SEO specialist to take a critical look at the filter structure of your webshop? We are happy to think along with you.
Conclusion
Faceted navigation is indispensable for the user and dangerous for your SEO if you let it run its course. The solution is not a single setting, but a deliberate choice per filter type: index what has search demand, canonicalise what is duplicate, block what is noise. Understand sharply what robots.txt, noindex and canonical each do and do not do, because that is exactly where the expensive mistakes arise.
Not sure how much crawl budget you are currently losing to filter combinations, or are the wrong URLs ending up in Google? Get in touch and we will review your filter structure in concrete terms.
Free website scan
Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.
We only use your details for your scan. No spam, unsubscribe anytime.