SEO & GEO
Block AI bots or let them in? The robots.txt decision for B2B
Copy for AI
For most B2B brands, blocking GPTBot in robots.txt is the wrong reflex. You then shut out precisely the bots that can surface your content in ChatGPT and Perplexity, while purchase-intent research increasingly happens there. In short, here you read the trade-off on whether to let AI bots in or block them on your site. Blocking only pays off for content you genuinely want to shield, such as paid reports or proprietary methodology. Below you read, per content type, when to keep bots out and when to let them in.
Check your own site: see in 10 seconds whether your robots.txt blocks AI bots with our free AI visibility check.
What does blocking AI bots actually do?
AI bots are the crawlers that models like ChatGPT and Claude use to read the web. The best known is GPTBot from OpenAI, alongside ClaudeBot (Anthropic), Google-Extended and Perplexity’s crawler, among others. In your robots.txt you can indicate per bot whether it may fetch your pages.
A block looks like this:
User-agent: GPTBot
Disallow: /
Important to grasp: this is an instruction, not a lock. It says “do not come here”, but it encrypts nothing. It only works because the big players respect the agreement. If you truly want to seal off content, it belongs behind a login or paywall, not merely behind a rule in robots.txt.
And mind the difference: keeping an AI bot out is not the same as keeping Google out. You can block GPTBot and keep ranking perfectly in the classic search results, because that is a different crawler (Googlebot). The choice therefore touches your AI visibility, not your classic SEO.
Why do B2B companies block GPTBot on reflex so often?
The reflex comes from a logical fear: “AI trains on my content and resells my knowledge without me getting anything out of it.” So the barrier goes down, often before anyone has even made the trade-off.
The problem is that this fear answers the wrong question. For an online shop with unique product photos, content theft is a real issue. For a B2B service provider that specifically wants to be found by the right decision-maker, invisibility is the bigger risk. Your knowledge articles, your approach, your cases: that is not a treasure you hide, it is your shop window.
And that shop window increasingly sits in an AI answer rather than in a list of blue links. A growing share of people now use AI tools alongside or instead of a traditional search engine, a shift that Google is also making around AI in search, and ChatGPT processes an estimated 2.5 billion prompts per day (TechCrunch). Part of those prompts are about suppliers, solutions and approaches in your market. If you block GPTBot, your brand simply cannot show up there as the answer.
That is also the broader shift behind the end of search: visibility moves from the search page to the generated answer. Whoever shuts off access puts themselves outside that answer.
When does blocking actually protect your revenue?
There are real cases where keeping bots out is the right choice, as analyses of who should block AI bots also emphasize. The common thread: block what undermines your revenue model if it ends up free in an AI answer, not what would actually help your visibility.
Blocking usually pays off for:
- Paid or premium content. Reports, whitepapers behind a form, paid knowledge bases. If the core of that appears without your brand in ChatGPT, you undermine your own offering.
- Proprietary methodology or IP. A unique framework, proprietary data or an approach that is literally your differentiator. What you protect commercially, you do not need to hand over freely for training.
- Customer portals and gated environments. Anything behind a login should not be reached by a public crawler in the first place. Here blocking is a hygiene measure, not a strategy.
- Sensitive or confidential pages. Internal documentation, customer-specific information, pricing you do not want to make public.
Note that all of this is content you would, in most cases, not simply put out publicly on your open site either. That is exactly the signal: if something belongs behind a wall, it also belongs out of reach of the bots.
When does blocking actually cost you leads?
For the lion’s share of a B2B site, the opposite holds. This is content you want as much reach as possible for, so you leave it open:
- Knowledge articles and blog. Your expertise is your best salesperson. If it gets cited in an AI answer, your brand enters at the moment someone is doing research.
- Service pages. When someone asks via ChatGPT “which agency helps with X in Belgium”, you want your offering to be able to feature in that answer.
- Cases and proof of results. Social proof works toward AI too: these are signals that help determine who gets proposed as the answer.
- About, team and authority signals. This feeds the picture the model builds of your entity.
The honest side of this story: granting access is in itself no guarantee. Letting a bot in does not automatically mean you get cited. Removing the block is a precondition, not a full strategy. If you want to actually show up in the answers, that calls for targeted GEO work. We help with that through our AI search optimization, where we always first look at whether the investment effectively leads to leads and not just to a nice number.
How do you decide per content type? A decision tree
Do not treat it as a single switch for your whole site. Walk through this for each content block:
- Is the content behind a login or paywall? Yes: block (and make sure the protection itself is correct). No: continue.
- Is it your proprietary IP or paid offering? Yes: block. No: continue.
- Do you want a prospect to find this during supplier selection? Yes: let it in. Doubt: let it in, because the default for public B2B content is visibility.
In practice that usually means: open for your public knowledge and offering pages, closed for specific gated paths. In robots.txt you handle that per path, for example by allowing everything but keeping one folder out:
User-agent: GPTBot
Disallow: /reports/
If you want a finer-grained consent layer toward AI crawlers, an llms.txt file is worth considering as a complement to robots.txt.
How do you set it up correctly without breaking anything?
A few practical points so you do not give away visibility or accidentally shut out Google:
- Use a separate rule per bot. GPTBot, ClaudeBot, Google-Extended and PerplexityBot are separate user-agents. A block for one does not touch the other.
- Do not confuse Google-Extended with Googlebot. Google-Extended steers AI training, Googlebot steers your classic search results. Never block Googlebot unless you truly want to vanish from Google.
- Test your file. A misplaced
Disallow: /accidentally shuts everything down. Check that public paths stay reachable. - Combine access with structure. Access plus good structured data helps models understand and cite your content correctly.
- Slow down instead of shutting out. If you do not want to shut a bot out but only limit its pace, a crawl delay in your robots.txt can temper the crawl speed, though not all crawlers respect that rule.
Do not do this as a one-off move. New AI crawlers keep appearing, your content grows, and the trade-off per content type shifts along with it. Schedule a half-yearly review of your robots.txt.
Frequently asked questions about blocking AI bots
Do I lose my Google rankings if I block GPTBot?
No. GPTBot and Googlebot are different crawlers. You can keep GPTBot out and simply keep ranking in the classic search results. Watch out with Google-Extended: that is Google’s AI-training variant, not your ordinary search crawler.
Is robots.txt real security against scraping?
No. It is a polite instruction that the big, well-behaved players respect. It does not encrypt or secure anything. Genuinely sensitive content belongs behind a login or paywall, not merely behind a rule in robots.txt.
My competitor blocks all AI bots, should I do that too?
Not without your own trade-off. If your competitor keeps itself out of AI answers and you do not, that is precisely where an opening arises for you. Decide based on your own content and goals, not on reflex.
Does granting access alone help you get cited?
Access is a precondition, not a guarantee. The bot has to be able to get in, but whether you actually appear in the answer depends on your authority, structure and relevance. That is what targeted GEO work is all about.
Ready to make the right choice?
The blocking question is not a technical detail, it determines whether your brand takes part in the AI answers where your buyers now choose. We look at it honestly with you: which content you better shield, which you open up, and whether investing in AI visibility actually delivers leads for you. No vanity numbers, but pipeline. Schedule your free intake.
Free website scan
Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.
We only use your details for your scan. No spam, unsubscribe anytime.