llms.txt: how to build it and its impact on AI visibility
llms.txt is a proposed text file meant to help language models quickly understand a website. What it contains, how to write it, and what we actually know about its impact on AI search visibility.
A website is built for a browser: layout, menus, scripts, ads. A language model trying to understand that same site has to filter out all that noise to extract useful information, with a limited context budget. llms.txt starts from this simple observation: offer, at the domain root, a summary written directly for a machine rather than for a browser.
The format was proposed in late 2024 by Jeremy Howard (answer.ai) and was quickly adopted by many technical documentation publishers. It is not (yet) an official standard the way robots.txt or sitemap.xml are, but Google has started checking for its presence in its Lighthouse audits dedicated to agentic browsing, which makes it a signal worth knowing about.
Key takeaways
- llms.txt is a Markdown file placed at the domain root (example.com/llms.txt), designed to be read by a language model rather than a browser.
- It lists, in plain language, what the site is, what it offers, and the most important pages or resources to check.
- It is not an official standard enforced by AI engines: no major provider (OpenAI, Google, Anthropic, Perplexity) currently confirms using it systematically for crawling.
- Google nonetheless includes it as one of the criteria in its experimental Lighthouse "agentic browsing" category, making it a readiness signal to watch rather than a guaranteed ranking lever.
What is the llms.txt file?
llms.txt is a Markdown text file, placed at a site's root (at /llms.txt), that summarizes the site's content in structured natural language. The idea is close to that of a sitemap.xml, but oriented toward meaning rather than technical exhaustiveness: instead of listing every URL, it highlights the pages that actually matter and explains in one sentence what each one contains.
Some sites also publish an llms-full.txt, a more complete version that includes the full text content of key pages rather than simple links, designed to be injected directly into a model’s context.
How to build an llms.txt file
The convention proposed by the format follows a simple, deliberately minimalist Markdown structure:
- An H1 heading with the site or product name.
- A blockquote (`>`) summarizing in one or two sentences what the site is and who it is for.
- Short paragraphs providing the context needed to understand the business, without marketing jargon.
- H2 sections (for example `## Documentation`, `## Guides`, `## Optional`) containing lists of Markdown links, each with a short description in parentheses.
llms.txt, sitemap.xml and robots.txt: three different roles
These three files coexist at a domain root but answer different questions. robots.txt indicates what a bot is allowed to crawl. sitemap.xml exhaustively lists the URLs to index, for a classic search engine. llms.txt replaces neither: it handles no permissions and does not aim for exhaustiveness. Its role is to save a model time by pointing it directly to the site’s most representative resources, with the context needed to interpret them correctly.
In practice, a site can perfectly well keep its classic sitemap.xml for traditional SEO while adding an llms.txt specifically designed for the AI layer.
What real impact on AI search visibility
It's worth being honest about where things stand: as of today, none of the major answer engines (ChatGPT, Gemini, Perplexity, Claude) has publicly confirmed systematically using llms.txt for crawling or grounding. It is not a guaranteed shortcut to more citations.
What changes the picture is that Google has started checking for this file’s presence in its experimental Lighthouse category dedicated to agentic browsing, alongside audits on WebMCP and machine accessibility. llms.txt thus becomes a technical readiness signal rather than a direct ranking lever: cheap to set up, it mainly forces a site to clearly state what it is and what matters in it, an exercise that also benefits the general content structuring useful for grounding.
Read next
llms.txt
The short definition of the term and its link to a site’s technical readiness for AI.
Agentic browsing: what Google recommends for AI agents
llms.txt is one of the Lighthouse audits dedicated to preparing a site for AI agents.
How AI engines select and cite their sources
The grounding mechanism, to understand what actually makes a page citable.
Methodology
Understand how AIglebot observes a brand’s presence in the answers of the tracked AI engines.
Generative Engine Optimization
Connect this technical readiness to the broader GEO approach to brand visibility.
Frequently asked questions
Should I really create an llms.txt for my site?
It is not essential, but it is cheap to set up and forces you to clarify what the site actually offers. It is also now one of the criteria Google checks in its agentic browsing audits, making it a useful readiness signal over the medium term.
Do ChatGPT, Perplexity or Gemini read this file?
None of these engines currently confirms systematically using llms.txt for crawling or answers. The format remains a community convention, voluntarily adopted by some tools and agents, with no guarantee of universal use.
What is the difference with sitemap.xml?
sitemap.xml exhaustively lists a site’s URLs for indexing by a classic search engine. llms.txt selects the most representative pages and describes them in natural language, to help a model quickly understand the site rather than index every URL.
Does llms.txt directly improve ranking in AI answers?
There is no evidence of that today. Its value is mostly indirect: it pushes you to clearly state the site’s offering and prioritize its key pages, which benefits the general content structuring useful for grounding.
AIglebot helps companies track their visibility, reputation and recommendation in ChatGPT, Gemini, Perplexity and Claude.