This is written for whoever has "add llms.txt" sitting on a roadmap, probably because an agency proposed it or an audit tool flagged it red.

Stop reading if you already know the answer is no and you just need something to forward. Send the evidence section below. It is the whole argument.

Keep reading if you want the tiered verdict, the format spec, and the honest case for the one situation where the file genuinely earns its place.

The short answer

For AI search visibility, no, you do not need llms.txt, and publishing one will not improve how often ChatGPT, Perplexity, Gemini or Google's AI Mode cite you. The evidence here is unusually clear for a search topic, which is rare enough to be worth saying plainly.

For agent and documentation use, it depends on what you publish. If you have developer documentation, an API, or a product an engineer integrates with, the file has a real consumer today and is worth building properly.

Here is the tiered version.

Your situationVerdict
Developer docs or an API productShip it, and ship llms-full.txt too. Coding agents fetch these today. The one case with observed demand.
B2B SaaS or services marketing siteShip a small one if it takes thirty minutes. Count it as zero. Option value on a cheap bet, not a tactic.
Content site or publisherSkip it, or ship and forget it. Your extraction problem is on-page, not at the root.
EcommerceSkip it. Product feeds and structured data do the job the file is imagined to do.
You were quoted for it as a deliverableAsk what the measurement plan is. There is not one that survives contact with the data below.

The nuance worth holding: harmless is not the same as useful. The cost of llms.txt was never the file. It is the attention it borrows from work that measurably moves citation share, and the false sense that the AI visibility box has been ticked.

What llms.txt actually is

llms.txt is a proposed Markdown file at the root of a domain that gives AI systems a curated map of a site's most important content. Jeremy Howard published the proposal on 3 September 2024, and llmstxt.org now carries a v2 revision dated 10 August 2026, informed by two years of adoption.

The reasoning was sound. Web pages are cluttered with navigation, scripts, cookie banners and footer junk, all of which burns model context before the useful content is reached. A clean Markdown index solves that in theory.

The spec is deliberately small. The structure, in order:

  • An H1 with the project or site name. This is the only required element in the entire spec.
  • An optional blockquote summarising what the site is, in a couple of sentences.
  • Optional free-form Markdown for context, using any element except headings.
  • Optional H2 sections containing link lists, formatted as a bulleted [name](url): note.
  • An optional ## Optional section at the end, for links an agent can skip when a shorter context is needed.

v2 also allows files at subpaths, so /docs/llms.txt can cover documentation while a root file covers the rest, with agents expected to use the most specific match.

There is a companion convention, llms-full.txt, which concatenates the full Markdown body of the listed pages into one file separated by dividers, so an agent can fetch everything in a single request. It is worth being precise here: llms-full.txt is a community convention, not part of the spec.

None of this is a standard. It is a proposal with a website. That distinction is the reason for everything below.

The evidence, in one place

Ahrefs, June 2026. They took all 137,210 domains in Ahrefs Web Analytics that received traffic in May 2026, checked each root for an llms.txt returning HTTP 200, then used bot analytics to examine every request to those paths, split by response code and classified by user agent.

The headline: 97% of valid files received zero requests during the month. Of 38,360 domains with a valid file, roughly 1,100 got any traffic at all.

Bar chart of who requests llms.txt files: SEO audit tools 21.7% at the top, AI retrieval bots at 1.1% near the bottom, alongside the finding that 97% of files received zero requests

The breakdown of who did fetch them is more damning than the headline. Of the requests that happened, SEO audit tools accounted for 21.7%, unidentified bots 14.9%, general web crawlers such as Googlebot 13.1%, and technology profiling tools such as BuiltWith 11.6%. AI bots across all categories made up 19.5%. AI retrieval bots specifically, the ones connected to ChatGPT and Perplexity that determine whether you appear in a live answer, made up 1.1%. GPTBot was 4.51% and ClaudeBot 0.8%. There is even a category for llms.txt discoverability bots, at 3.6%.

Every one of those percentages is a share of the 3% of files that received any request at all. Keep that qualifier attached whenever you quote them.

The negative result inside the same study. Ahrefs also looked at requests for llms.txt on domains where the file does not exist. The AI bot share of those 404s was zero, and 98% of that traffic was human, most plausibly SEOs manually checking competitors.

That is the detail that settles the argument. A crawler that wanted the file would ask and be refused. These crawlers never ask.

SE Ranking, roughly 300,000 domains. No significant correlation between having an llms.txt file and how often a domain is cited in AI answers. They went further than a correlation test and built an XGBoost model of citation drivers: removing llms.txt as a feature made the model more accurate. The file was adding noise, not signal.

Meanwhile publication keeps climbing. Ahrefs found 28% of its analytics customers publish one, and flagged that as an upper bound because its customer base skews technical and SEO-aware. SE Ranking's broader crawl found 10.13%. Both numbers are real and the gap between them is just sample composition, which is worth knowing because the two get quoted against each other as if one is wrong.

So publication is accelerating and consumption is not. That gap is the entire story of llms.txt in 2026.

The vendor statements line up with the logs. Gary Illyes confirmed at Search Central Live APAC in July 2025 that Google does not support llms.txt and has no plans to. John Mueller went further, comparing it to the keywords meta tag: a self-declared description of what a site owner claims their site is about, which is exactly the kind of signal that is too easy to game to be trusted.

One clarification worth making, because the counter-argument gets used a lot. llmstxt.org points at OpenAI, Anthropic and Google as adopters, and that is true. They publish llms.txt files for their own documentation. None of them has committed to reading yours during retrieval. Publishing is not consuming, and the fact that every prominent adopter is a documentation site is not a coincidence. It is this post's thesis arriving early.

Why it does not work: the handshake problem

The mechanics are worth understanding, because they predict what will happen to the next file someone proposes.

Two panels comparing robots.txt, where publishers write the file and crawler operators agreed to read it, with llms.txt, where only the publisher half was ever shipped

robots.txt works because crawler operators agreed to read it. The convention is one half of a handshake and the other half was accepted. llms.txt shipped the publisher half of a handshake that no model provider ever agreed to. Publishing a file does not create an obligation to read it, and thousands of sites publishing it does not either.

There is a second reason, and it is the one Mueller's comparison points at. An assistant answering a question already has your pages, from a search index or a live fetch. A self-authored file claiming which of your pages matter adds nothing the system trusts more than the pages themselves, because the file is exactly the kind of source with an incentive to overstate. Unverified self-description has lost this argument before, repeatedly, since roughly 1997.

This is why "but they might adopt it later" is weaker than it sounds. The reason to adopt it later would be that it solves a retrieval problem. Retrieval is not currently constrained by not knowing which pages a site owner considers important. It is constrained by whether the content is good, extractable and corroborated elsewhere.

Google's two voices, and how to reconcile them

In May 2026 Google managed to take both sides of this within about a week, which produced most of the confusion currently in circulation.

On 7 May, Lighthouse 13.3 moved its Agentic Browsing category from experimental into the default configuration, in Chrome 150 and later. The category evaluates how well a site is built for machine interaction using four deterministic pass or fail audits rather than a weighted score, because the standards are still forming. Among the checks: WebMCP integration, accessibility tree integrity, layout stability, and the presence of a retrievable llms.txt.

Days later, Google Search Central's guidance on optimising for generative AI features told site owners, in as many words, that they do not need to create new machine-readable files, AI text files or markup to appear in those features.

Both are correct, because they govern different things.

Search's guidance covers discovery: crawl, index and source selection, the process that decides whether you appear in AI Overviews and AI Mode. On that axis llms.txt does nothing.

Lighthouse's category covers post-landing interaction: whether an agent already on your site can orient itself and operate it. Different interaction model, different requirements, and not a ranking system.

The one consumer that is real

Here is the part most of the critical coverage skips, and the reason we do not tell every client to delete the file.

In the Ahrefs breakdown of AI fetchers, GPTBot was top and Claude-Code came next, exceeding every AI retrieval bot, assistant and training crawler. That is a coding agent, not a search crawler. It tells you what the file is actually being used for by the small number of systems using it: an agent orienting itself in a codebase or a documentation set, deciding what to read next.

That maps to who publishes the well-known examples. Anthropic, GitHub, Cloudflare, Mintlify. Documentation-heavy technical products, where a developer or their agent needs to find the right reference page quickly and the marketing site is beside the point.

So the honest scoping is this. llms.txt is not an AI search visibility file. It is a documentation navigation file that got marketed as an AI search visibility file, and it works reasonably well at the first job.

If you publish an API, an SDK, an integration guide, or anything an engineer implements against, build one and build llms-full.txt alongside it. The consumers exist and the file saves them work.

If you publish a services site with fifteen pages and a blog, the file has no job to do. Ship it in half an hour because it costs half an hour, and move on.

So what actually goes in it

If you are shipping one, ship it correctly. Most published files are junk, which is part of why nobody reads them.

The H1. Your company or product name. Nothing else. It is the only required element in the spec.

The blockquote. The highest-leverage line in the file and the one most often skipped. One or two sentences defining what you are, factually. Systems that do read the file treat this as the canonical one-line definition of your entity, so it should read like a reference entry rather than a positioning statement. "Omnitics is an AI demand generation agency for B2B and B2B2C companies" is useful. "Omnitics is the leading partner for ambitious brands" is noise, and superlatives specifically signal marketing copy to a system trying to extract facts.

Context paragraph, optional. A few lines of free-form Markdown, no headings, explaining how to interpret the sections that follow. Useful when your site structure is non-obvious. Skip it if it is not.

H2 sections with link lists. Group by function, not by your navigation. Name the sections specifically: "API Reference" and "Implementation Guides" beat "Docs" and "Resources". Each entry is a bulleted Markdown link, then a colon, then a single sentence describing what the page covers and when it matters.

The Optional section. Anything a model can skip when context is tight: changelogs, blog archives, press. A genuinely useful part of the spec that almost nobody uses, ourselves included.

Constraints that matter. Absolute HTTPS URLs, never relative paths. Plain Markdown only, no HTML or JSON-LD. Serve it as text/markdown if your CDN allows, though text/plain works. Keep the root file small, because the entire point is selection rather than completeness. Your sitemap is for completeness. If a page does not earn its bullet, cut it.

llms-full.txt, only if you have docs. Concatenated full Markdown of the listed pages, divider-separated, at /llms-full.txt. Genuinely useful for single-fetch agents. Genuinely pointless for a twelve-page marketing site, where it becomes a maintenance liability with no consumer.

A working example

For a B2B services site the whole thing should look roughly like this, and take under thirty minutes. This is a trimmed version of the file we publish at omniticshq.com/llms.txt.

# Omnitics

> AI demand generation agency for B2B and B2B2C companies. Programmatic ABM,
> hyper-personalised cold email, webinar and podcast flywheels, SEO, AEO and
> GEO, agentic n8n workflows, and HubSpot and Salesforce operations.

Sections below are grouped by what the page is for. Service pages describe
delivery scope. Playbooks are long-form methodology posts.

## Services

- [Account-Based Marketing](https://omniticshq.com/services/abm): Target account
  selection, tiering, plays and measurement for B2B teams.
- [SEO, AEO and GEO](https://omniticshq.com/services/seo-aeo-geo): Organic and
  AI-answer visibility, including citation tracking.

## Playbooks

- [AEO: Winning the Answer Box, People Also Ask, and Voice](https://omniticshq.com/blog/aeo-answer-box-people-also-ask-voice):
  How answer extraction works across search surfaces in 2026.
- [GEO Explained](https://omniticshq.com/blog/generative-engine-optimization-guide):
  Framework for getting a brand cited by AI assistants.

## Company

- [About](https://omniticshq.com/about): Team, background and how we work.
- [Contact](https://omniticshq.com/contact): Enquiries and strategy calls.

## Optional

- [Blog archive](https://omniticshq.com/blog/): Full post index.

That is the correct shape. Notice how little there is to get wrong, which is both the appeal of the format and the reason it cannot carry the strategic weight assigned to it.

The four mistakes that cost you something

The file is harmless. Some implementations of it are not.

  1. Publishing indexable Markdown copies of every page. The most damaging pattern, and it circulates in a lot of tutorials. Generating a .md twin for each URL and leaving those files indexable creates duplicate content at scale, which dilutes crawl budget and can suppress the originals. Since conventional search authority remains a primary input into whether AI systems treat you as credible, hurting your SEO to chase an AI signal that does not exist is a strictly negative trade. If you generate Markdown twins, noindex them.
  2. Letting it go stale. A file listing pages that 404, or describing a product you renamed, is worse than no file, because the small number of systems that do read it now hold wrong information about you. Put it on the same review cycle as your sitemap, or generate it.
  3. Treating it as a blocking mechanism. It is not one. llms.txt cannot restrict any crawler or prevent any system from reading your site. Access control lives in robots.txt, at your CDN, or in commercial terms.
  4. Reporting it as a completed GEO milestone. The real cost. A programme that ships llms.txt in month one and reports AI visibility as underway has spent its credibility on a file that 97% of the time is never requested. When the citation numbers do not move in month four, the whole programme absorbs the blame.

Our own file is generated by the CMS from the same database that renders the site, which is the cheapest possible answer to mistake number two. It cannot list a page that no longer exists, because it is rebuilt whenever the site is. If you are going to publish one, generate it rather than hand-maintaining it.

What to do with the same hour instead

Ranked by observed effect on citation share, highest first. None of these are exotic.

  1. Fix the answer structure on pages that already rank. Self-contained sections, the answer in the first forty to fifty-five words, one checkable fact per section. This is the highest-yield hour available to most B2B sites, and we set out the full spec in the AEO playbook.
  2. Earn third-party mentions in the places that get cited. Roundups, directories, review platforms, community threads. Corroboration from sources you do not control is what separates being extractable from being recommended, which is the GEO half of the job.
  3. Publish one piece of original data. A number nobody else has, that other people quote, is the most durable citation asset there is. It also solves the problem that most of your content cites other people's research.
  4. Get your entity right. Organization schema with sameAs, consistent naming across every profile, a real author with a real professional presence. Machines need to know you are a distinct thing before they can recommend you.
  5. Check what your robots.txt is actually doing to AI crawlers. Most sites have never audited this. Providers now separate training crawlers from search crawlers, so you can block GPTBot and Google-Extended while allowing OAI-SearchBot and PerplexityBot, which are the ones that make you eligible for citation. A surprising number of sites are accidentally blocking the second group.
  6. Then, if you like, ship llms.txt. Thirty minutes, no expectations.

Number five is the one that catches people out. Blocking a training crawler is a defensible commercial decision. Blocking a retrieval crawler removes you from live answers entirely, and the two get confused constantly because the naming is inconsistent across providers.

WebMCP is the thing to actually watch

The second item in Lighthouse's Agentic Browsing audit is WebMCP, and it is a considerably bigger deal than the file this post is about.

llms.txt as a static description an agent reads, next to WebMCP as a live interface an agent calls to operate a site

WebMCP lets a site declare structured tool contracts, in HTML attributes or JavaScript, so an agent can act on a live session directly instead of scraping the DOM or driving your interface with screenshots. It is being developed in the W3C Web Machine Learning Community Group, with the API contributed jointly by Google and Microsoft, and it entered a public origin trial in Chrome 149 announced at the Google I/O developer keynote on 19 May 2026.

The difference in ambition is the point. llms.txt is a static description an agent reads to orient itself. WebMCP is a live interface an agent calls to do something. One describes a site, the other lets an agent operate it. If the agentic web develops the way Chrome is provisioning for, an action layer beats a description file, and the same logic that made llms.txt unnecessary for retrieval makes a structured action layer necessary for transactions.

For B2B specifically, this connects to something we covered when 6sense and Demandbase both shipped MCP servers earlier this year. The pattern is consistent across the stack: vendors are concluding that the interface is not where the work happens, and are exposing callable structure instead. Your website is on the same trajectory, just earlier.

Nothing to implement urgently unless you have transactional flows worth exposing. Everything to watch.

Where we land

llms.txt is a good idea that solved a real problem for a hypothetical consumer who never showed up. The proposal was thoughtful, the format is sensible, and the one population using it, coding agents reading documentation, is being served exactly as intended. Everyone else published a file into a void and called it AI optimisation.

Ship one if you have documentation, because there is genuine demand. Ship one anyway if thirty minutes is genuinely free, because option value is real and the downside is nil. Never let it appear on a roadmap as a visibility tactic, never pay for it as a deliverable, and never report it as progress.

The broader lesson is worth more than the file. When a new AI-era convention appears, the question is not whether it exists or whether it is easy. The question is whether the consumer side of the handshake ever agreed. Check the server logs before you check the tutorials. Everything llms.txt needed to tell us was sitting in request data long before the industry finished arguing about it.

At Omnitics we run AI visibility off measured citation share rather than checklist completion, which occasionally means telling clients that the deliverable they asked for is not worth building. That is what our SEO, AEO and GEO practice actually does.

Want to know what your AI files are actually doing?

Bring your domain to a 30-minute call. We will look at what your llms.txt and robots.txt are actually doing, show you which AI crawlers are reaching you and which are being blocked by accident, and tell you where the hour is better spent. If the answer is that the file is doing nothing, we will say so.

Book your strategy call

Frequently asked questions

llms.txt is a proposed Markdown file at a website's root that gives AI systems a curated map of the site's most important content. Jeremy Howard published the proposal on 3 September 2024, with a v2 revision in August 2026. It requires only an H1 with the site name, and optionally includes a summary blockquote and H2 sections listing links with one-line descriptions.

Overwhelmingly, no. Ahrefs analysed 137,210 domains and found 97% of llms.txt files received zero requests during May 2026. Of the requests that did occur, AI retrieval bots accounted for 1.1% while SEO audit tools accounted for 21.7%. On domains with no such file, the AI bot share of those 404 requests was zero, so the crawlers are not probing for it either.

No. Gary Illyes confirmed in July 2025 that Google does not support llms.txt and has no plans to, and Google's 2026 AI features guidance says you do not need to create new machine-readable files or AI text files to appear in its generative features. SE Ranking found no significant correlation between the file and AI citation frequency across roughly 300,000 domains.

Ship one if you publish developer documentation or an API, where coding agents genuinely fetch it. For a standard marketing site, ship one only if it takes about thirty minutes, and treat it as a cheap option on future adoption rather than a visibility tactic. Do not pay an agency for it as a deliverable.

An H1 with your company name, a factual blockquote defining what you do in one or two sentences, then H2 sections grouping your most important pages as Markdown links with a one-sentence description each. Use absolute HTTPS URLs, keep the root file small because the point is selection rather than completeness, and put low-priority content under an Optional heading at the end.

llms.txt is an index of curated links with descriptions. llms-full.txt concatenates the complete Markdown content of those pages into a single file so an agent can retrieve everything in one fetch. llms-full.txt is a community convention rather than part of the spec, is worth building for documentation sites, and is generally a maintenance burden without a consumer for small marketing sites.

No, and the comparison is misleading in a way that matters. robots.txt controls crawler access and is honoured by every major crawler because those operators agreed to read it. llms.txt describes content and is honoured by essentially none. llms.txt cannot block, restrict or prevent any AI system from reading your site.

The two teams are measuring different things. Google Search's guidance covers discovery and citation, where the file has no effect. Lighthouse 13.3's Agentic Browsing category, which moved into the default configuration on 7 May 2026, covers whether a browser-based agent already on your site can orient itself and operate it. That category runs deterministic pass or fail audits, carries no 0 to 100 score, and does not influence search results.

The file itself is harmless. Two common implementations are not. Publishing indexable Markdown duplicates of every page creates duplicate content at scale, which dilutes crawl budget and can suppress your original pages. Letting the file go stale means the few systems that do read it hold outdated information about you.

Restructure pages that already rank so the answer is extractable, earn third-party mentions on sources AI systems cite, publish original data others quote, get your Organization and author markup right, and audit your robots.txt to confirm you are not accidentally blocking retrieval crawlers such as OAI-SearchBot and PerplexityBot while intending to block training crawlers.

Possibly, and the cost of hedging is thirty minutes. But adoption would require the file to solve a retrieval problem that does not currently exist, because assistants already have your pages and have historically discounted unverified self-description. Treat future adoption as a cheap option rather than a forecast.

Sid R
Sid R · GTM & Demand GenWorked with companies like CleverTap, Sprinto, Netcore and have been an Ex-founder. Overall has 17 strong years of Growth Marketing Experience. Book a strategy call.View LinkedIn