· 7 min read · Wwwebtech Team

llms.txt: Worth Ten Minutes, Not a Retainer

What llms.txt actually is, who reads it today, which robots.txt controls genuinely work, and a sane position to take on AI access this year.

Somebody has probably told you your site needs an llms.txt file. Perhaps a plugin offered to generate one. Perhaps an agency put it in a proposal under a heading like "AI readiness". Before you pay for it, it is worth understanding what the file is, who currently reads it, and what the honest answer is to "will this get my business into ChatGPT's answers".

The short version: creating the file is cheap and harmless, the claims made for it are mostly unproven, and the controls that genuinely do something today sit somewhere else entirely — in your robots.txt file, which has existed since the 1990s.

What llms.txt actually is

In September 2024, Jeremy Howard of Answer.AI published a proposal for a file called llms.txt, placed at the root of a website — so yoursite.in/llms.txt, in the same spot as robots.txt.

The idea is reasonable. A large language model — the technology behind ChatGPT, Gemini, Claude and Perplexity — reads text. It works with a limited amount of text at a time. When one of these tools fetches a page of yours, it gets the whole HTML document: navigation, cookie banner, footer, three scripts, a chat widget, and somewhere in the middle, the 300 words that actually answer the question. The proposal says: give the machines a clean, curated map instead.

The format is Markdown, which is just plain text with light formatting. A valid file looks roughly like this:

  • A single top-level heading with your business or site name.
  • A short summary line explaining what the site is.
  • Sections of links, each with a one-line description — your services, your key guides, your contact page.
  • Optionally, a second file called llms-full.txt containing the actual content in plain text rather than just links.

That is the entire specification. It is not complicated. A competent developer can write one for a twenty-page site in under an hour, and most of that hour is deciding what to include.

Who actually reads it today

Here is where the sales pitch and the evidence part company.

As things stand, no major AI company has publicly documented that its crawlers look for llms.txt or treat it as a signal. OpenAI, Anthropic, Google and Perplexity all publish documentation about how their crawlers behave — which user-agent names they send, how they respect robots.txt. None of that documentation tells you to write an llms.txt file. Public comments from Google staff have been openly sceptical, with comparisons drawn to the old keywords meta tag, which search engines stopped using decades ago because it described what site owners wanted to rank for rather than what the page said.

You can test this yourself, and you should. If you have an llms.txt file, look in your server access logs for requests to that path, and note which user-agents made them. We have written before about reading logs for AI crawler activity. In most logs the file is either never requested, or requested by tools that are checking whether it exists rather than reading it to answer a question.

That may change. It may change quickly. But "a sensible idea that a vendor might adopt" is a different product from "a control that works today", and you are entitled to be told which one you are buying.

The controls that genuinely work

While llms.txt is a proposal, robots.txt is a real, long-standing convention, formalised as RFC 9309 in 2022, and the major AI companies document their compliance with it. Each publishes the user-agent token its crawler sends, so you can allow or disallow them by name.

The crucial distinction — the one most owners miss — is between crawlers that gather text to train future models and fetchers that retrieve a page right now to answer a live question and cite it. These are usually different tokens, and blocking the wrong one has the wrong effect.

Token in robots.txtRoughly what it does
GPTBotOpenAI's bulk crawler, associated with model training
OAI-SearchBotOpenAI's crawler for its search and citation index
ChatGPT-UserFetches a page because a user asked about it in that moment
ClaudeBotAnthropic's bulk crawler
PerplexityBotPerplexity's indexing crawler
Google-ExtendedNot a crawler. A permission flag covering Gemini and grounding
CCBotCommon Crawl, a public dataset many models have been trained on

Two consequences follow directly. If you block GPTBot to keep your writing out of training data, you have not removed yourself from ChatGPT's cited answers — that is a separate token. And Google-Extended does not control whether you appear in AI Overviews inside Google Search; those are served off the ordinary Search index, and the ordinary Search controls apply, including nosnippet and max-snippet, which also reduce your normal result snippet. There is no switch that says "appear in blue links, vanish from the AI summary above them".

Separately, the IETF — the standards body that produces internet protocols — has an active working group on AI preference signalling. That work is genuinely in progress and may eventually produce something more expressive than a list of bot names. It has not produced a finished standard you can deploy today.

What I would not pay for

Three things, all currently sold.

A monthly retainer whose main deliverable is an llms.txt file. It is a static text file. If the work is "we will keep your llms.txt updated", ask what changes month to month on a site that publishes twice a quarter. This is a configuration task, not a service line.

Auto-generated files that dump every URL on the site. A plugin that lists 400 pages, including tag archives and paginated listings, with no descriptions, has produced a worse sitemap in a format nothing reads. The entire value of the proposal is curation. If nobody chose what goes in, nothing of value was created.

Any promise that this gets you quoted. Nobody can promise that, because the people who run these systems have not said how they choose sources. What is observable is that AI tools tend to lift passages that answer a question cleanly in the page's own words. That is a content problem, not a file-format problem — our view on AI visibility work is that the writing does the heavy lifting and the plumbing merely stops you getting in your own way.

A defensible position to take now

Here is what I would actually do, in order of how much it matters.

  1. Decide your policy on AI access deliberately. Not by default, not by plugin. If you are a publisher whose text is the product, blocking training crawlers is defensible. If you are a dentist in Preet Vihar who wants to be recommended, blocking AI retrieval bots is self-harm. Most Indian SMEs are in the second group.
  2. Write that policy into robots.txt by token name, keeping the distinction between training and retrieval bots clear. Check it after any site migration — a rebuild very often ships with a staging robots.txt that blocks everything.
  3. Make your important pages answerable. Clear headings, a direct first sentence under each, prices or ranges if you can publish them, service areas named, hours and phone number in text rather than inside an image. This also happens to be ordinary good search work.
  4. Add structured data — the machine-readable markup Google documents for organisations, local businesses, products and FAQs. Unlike llms.txt, this is read today by systems you already care about.
  5. Then, if you like, add an llms.txt file. Hand-written, twenty lines, pointing at your ten pages that matter. Treat it as a cheap option on a future convention, not as a lever you pulled.

The reason to do step five at all is asymmetry. The cost is one hour, once. If the convention is adopted, you are already set up. If it is quietly abandoned — as plenty of well-intentioned web proposals have been — you have lost an hour and a file nobody requests. That is a reasonable bet. Paying ₹X every month for it is not.

Be honest with yourself about the uncertainty here. Anyone telling you confidently how AI systems pick their sources is guessing, including people with large followings. What can be verified is which bots hit your server, what your robots.txt says, and whether your pages actually answer the questions people ask. Work on those and you are covered under most futures.

What to do next

Open yoursite.in/robots.txt in your browser right now and read it. If you do not understand a line of it, that is the problem to fix first — not the file you have not got. Then ask whoever hosts your site to pull a week of access logs and tell you which AI user-agents appeared and which paths they fetched. That ten-minute check tells you more than any article will.

If you want a second pair of eyes on your crawler policy, your structured data and whether your key pages read like answers, tell us what you are trying to achieve and we will look at what you have. If your site is already falling over under bot traffic, that is a different conversation and belongs with technical support.

Questions we get asked

Do I need an llms.txt file for my website?

You do not need one. No major AI company has publicly documented that it reads the file, so nothing today depends on it. It costs about an hour to hand-write a short one, which is a cheap bet in case the convention is adopted, but it should sit at the bottom of your list, below clear page content and correct robots.txt rules.

Will blocking GPTBot stop my business appearing in ChatGPT?

Not necessarily. OpenAI documents separate crawler names for different purposes, and the bot associated with bulk training is not the same as the ones that fetch pages to answer a live question and cite them. Blocking one and not the others produces a mixed result, so decide deliberately which behaviour you actually object to.

Can I appear in Google's blue links but not in AI Overviews?

There is no clean switch for this. AI Overviews are served from Google's ordinary Search index, and the available controls — such as nosnippet and max-snippet — also shrink or remove your normal result snippet. You are generally choosing between being summarised and being less visible overall.

Is llms.txt the same thing as a sitemap?

No. An XML sitemap is a machine-readable list of URLs that search engines genuinely use for discovery. llms.txt is a proposed Markdown file with a curated handful of links plus plain-English descriptions, intended as a summary for language models. A file that simply lists every page on your site has recreated a worse sitemap in a format nothing currently reads.

How do I tell whether AI crawlers are visiting my site at all?

Ask your host or developer for your server access logs and filter by user-agent for names like GPTBot, ClaudeBot, PerplexityBot and OAI-SearchBot. The logs show the date, the path requested and the response code, which tells you both who came and whether they got a working page or an error.

If this is your problem

What we’d actually do about it.

All posts

Start here

Want us to look at yours?

Send the URL and what you think is wrong. We’ll tell you what we see, whether or not you hire us. Reply within 1 business day.

What do you need?

We reply within 1 business day. No newsletter, no sales sequence.