Customer Impact

SEO & GEO

What is llms.txt and do you need one? How to decide who trains your content

Copy for AI

llms.txt is a simple text file you place on your website to neatly point AI systems toward your most important content. The short version: it is not a ranking factor today and it enforces nothing, but it belongs to a broader question every B2B brand has to answer: do you let AI crawlers in (more chance of visibility in AI answers) or do you keep them out (protecting your content)? For most B2B companies, allowing more is the right call, and below we explain honestly when that is different.

Measure it yourself: see how ready your page is to be cited by AI with the free GEO check.

What exactly is llms.txt?

llms.txt is a markdown file in the root of your website (so yourdomain.be/llms.txt) that gives AI models a tidy overview of your most important pages. Think of it as a kind of menu card for language models: instead of a model having to plough through your entire site with its menus, footers and cookie banners, it gets a clean list of links and short descriptions of what really matters.

The idea leans on robots.txt and sitemap.xml by name, but the goal is different, as the original explanation of the llms.txt proposal also emphasises. A sitemap says “here are all my pages”. robots.txt says “you may or may not crawl these”. llms.txt mainly says “this is the content that matters to you, and this is how you should understand it”. So it is about readability and context, not about enforcing access.

Important to set straight right away: llms.txt is a voluntary convention. No AI provider is obliged to read or respect it. It is a proposal, not a law, and so far the major models are not demonstrably adopting it.

Does llms.txt affect your SEO or AI visibility?

Here we have to be honest, because that is what we build our advice on. At this moment, llms.txt has no direct ranking impact. Search engines do not interpret it for SEO, and there is no evidence that an llms.txt file directly improves your positions in Google or your mentions in ChatGPT. What it does do is help determine how easily a model can understand your content and use it in generative answers.

In other words: it is a readability aid, not a miracle cure for visibility. Anyone who creates an llms.txt and thinks the leads will roll in by themselves will be disappointed. The real levers for AI visibility sit elsewhere: in the five core indicators that determine whether a model recognises your brand, understands it correctly and recommends it. llms.txt is at most a small, supporting part of that.

That is exactly our thesis at GEO: steer on visibility that delivers leads, not on a tick list of technical files. An llms.txt without a strong entity, consistent mentions and good content is a pretty facade on an empty house.

llms.txt versus robots.txt: what governs what?

This is where most of the confusion arises, so let us make it sharp. These two files solve different problems.

  • robots.txt governs access. This is where you allow or block crawlers, including AI crawlers such as GPTBot (OpenAI), ClaudeBot (Anthropic) and the Google-Extended directive. If you really want to prevent an AI system from retrieving your content to generate answers or train models, this is where you do it.
  • llms.txt governs readability. It improves how a model understands the content it is allowed to access, but it keeps nobody out.

The history helps to place this. OpenAI launched GPTBot in 2023 with an opt-out path for websites, after which Google followed with Google-Extended and Anthropic with its own crawler (OpenAI). Control over training and crawl access therefore runs through those crawler directives, not through llms.txt. Anyone who wants to shield their content needs robots.txt and those specific user agents. We wrote a separate, extensive guide about that: blocking AI bots via robots.txt.

When does blocking pay off, and when does it not?

This is the decision that really matters, and we would rather give honest advice than a one-size-fits-all rule. The trade-off is simple to frame: blocking protects content, allowing delivers visibility.

Allowing pays off for most B2B brands. Your buyers increasingly do their preliminary research in ChatGPT and Perplexity instead of in a classic search engine, a shift that Google itself is accelerating with AI in search. If your knowledge articles, product pages and expert content are not allowed to play there, you miss exactly the moment when a prospect compares suppliers. For a service provider that sells on authority, invisibility in AI answers is more expensive than any theoretical content risk.

Blocking pays off in specific cases. Do you genuinely have proprietary content, paid knowledge bases behind a login, or material that forms your competitive advantage? Then shielding it is defensible. That is a deliberate IP decision, not a reflex.

What we see too often: B2B companies that shut everything down out of fear and thereby cut off their own findability in exactly the channel where buying intent takes shape. Do not block on principle. Only block what you really must protect, and let the rest work for your visibility.

How do you create an llms.txt file?

If you decide an llms.txt fits, creating it is easy. You need no tooling or developer for it.

CREATING LLMS.TXT From empty file to llms.txt in the root 1 Create file llms.txt 2 H1 + intro brand name 3 Sections with links 4 Keep it clean up to date 5 In the root /llms.txt
Creating an llms.txt in five steps, without tooling or a developer.
  1. Create a markdown file named llms.txt.
  2. Start with an H1 carrying your brand name, followed by a short blockquote that sums up in one sentence who you are and who you serve.
  3. Add sections (for example “Services”, “Resources”, “About us”) with links to your most important pages underneath, each with a short description.
  4. Keep it clean and current. Limit yourself to the pages that really matter to anyone who wants to understand your brand.
  5. Place the file in the root of your domain, so it is reachable at yourdomain.be/llms.txt.

A minimal example:

# Your Company
> B2B growth partner for [audience] in Flanders and Brussels.

## Services
- [AI search / GEO](https://yourdomain.be/diensten/geo-agency/): visibility in AI answers.

## Resources
- [What is GEO](https://yourdomain.be/kennis/...): an explanation of generative engine optimization.

Do not count on this file to attract traffic by itself. See it as a neat support layer on top of the real work: a strong, consistent entity and content that AI models like to cite.

What would Customer Impact recommend?

For most B2B brands, our approach would sound like this: simply let AI crawlers into your public content, because that is where your visibility opportunity sits. Create an llms.txt if you wish, as a cheap and tidy layer, but expect no ranking miracle from it. Put your real energy where it pays off: making sure AI models understand your brand correctly and recommend it at the moment a buyer chooses a supplier.

That is measurable work that delivers leads, not a tick on a technical checklist. Do you want to know whether your brand shows up in AI answers today, and which interventions really make the difference for your pipeline? We are a small team that moves fast and says honestly when something is not worth it.

Book your free intake

Frequently asked questions about llms.txt

Do I need an llms.txt?

“Need” is a big word. It is not a ranking factor and your site works fine without one. It can be a neat, cheap addition that helps AI models understand your content, but it does not replace a real GEO strategy. Prioritise your visibility and authority first.

Does llms.txt improve my position in ChatGPT or Google?

Not directly. There is no direct ranking impact today: search engines do not interpret llms.txt for SEO. It does help determine how well a model can read the content it is allowed to access and use it in generative answers.

What is the difference between llms.txt and robots.txt?

robots.txt governs access (who may crawl your content, including AI bots). llms.txt governs readability (how a model understands the content it is allowed to access). If you want to block AI crawlers, you do that via robots.txt and crawler directives, not via llms.txt.

Should I block AI crawlers to protect my content?

Only if you genuinely have proprietary or paid content. For public B2B content, blocking mainly cuts off your own visibility, right where buyers do their research today. Block deliberately, not out of reflex.

How do I know whether AI represents my brand well?

By measuring it. Start with the five core indicators of AI visibility and look at how GEO compares to classic SEO, so you know which interventions really improve your mentions in AI answers.

Free website scan

Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.

Where should we send your report?

We only use your details for your scan. No spam, unsubscribe anytime.