La Boétie
Tools · one run · $25 or 20 €

llms.txt. The file that tells an AI what your business actually is.

Put your domain in. We read every page of your site, write your llms.txt, and audit your SEO and GEO while we are in there. Close the tab if you like: your file arrives by email.

  1. Enter your domain
  2. Pay $25
  3. We read your site
  4. Download your files

No account · Files by email · Refunded if it fails

GPTBot · ClaudeBot · OAI-SearchBot · Amazonbot · DeepSeekBot

We see them in our own server logs, week after week, on the sites we run. Your file is built to the same structure.

What your 25 dollars actually buys.

Six stages, run by an agent that reads your site the way a researcher would. It takes as long as your site is big, and you can close the tab through all of it.

  1. 01

    Discovery

    robots.txt, sitemap.xml and every nested index, or a link crawl when there is no sitemap. A full URL inventory, classified by kind.

  2. 02

    Deep read

    Your pages read in parallel, commercial substance first. Purpose, audience, entities, offers, canonicals, headings and existing structured data extracted from each.

  3. 03

    Synthesis

    What the business is, who it serves, what it sells, how it differs, what it can prove.

  4. 04

    Composition

    The file written section by section in your site's language, every URL verified live before it is committed.

  5. 05

    Audit

    A versioned SEO and GEO checklist scored over everything already gathered. Coded, severity ranked, with evidence.

  6. 06

    Delivery

    Downloadable as .md or .txt on screen, and emailed so you keep a permanent copy.

These are the crawlers that come for it.

We see them in our own server logs on the sites we run. Every one of them writes answers that people read instead of clicking through to you.

  • OpenAIGPTBot, OAI-SearchBot, ChatGPT-UserChatGPT and ChatGPT Search
  • AnthropicClaudeBot, Claude-UserClaude and Claude Code
  • GoogleGoogle-CloudVertexBotGemini and Vertex AI
  • AmazonAmazonbotAlexa and Rufus
  • DeepSeekDeepSeekBotDeepSeek
One run

$25, or 20 €. Once.

No subscription, no account, no call. Give us a URL and an email address, pay, and close the tab if you want. The run keeps going without you and the files are waiting when you come back.

  • A researched llms.txt, written from your actual pages
  • Every URL in it verified live before it is committed
  • As .md and .txt, downloadable and emailed to you
  • A full SEO and GEO audit of your site, run at the same time
Generate mine
Refunded if we cannot deliver
FAQ

Questions before you start.

A plain text file in Markdown at the root of your domain, at https://yourdomain.com/llms.txt, that tells a large language model who you are, what you offer and where the substance on your site lives. It was proposed by Jeremy Howard of Answer.AI in September 2024 and adoption has grown fivefold in the last twelve months.

The complete guide to llms.txt

Everything we know about the format, which engines read it, how a model decides to cite you, and what separates a file that earns a fetch from one that does not.

What llms.txt actually is

A plain text file at your domain root, written for the AI that has to summarise you in one read.

Jeremy Howard of Answer.AI published the llms.txt proposal on 3 September 2024. The idea is narrow and sensible. A large language model reading your website has a context budget, and your website spends that budget badly. It spends it on navigation menus, cookie banners, footer link farms, tracking scripts, and the same header repeated on nine hundred pages. By the time the model reaches the sentence that explains what you sell, the budget is gone.

llms.txt is the correction. One file, at https://yourdomain.com/llms.txt, written in Markdown, that states who you are, what you offer, and where the substance lives. No markup to parse, no JavaScript to execute, no crawl to run. The convention adds a second optional file, llms-full.txt, which inlines the actual content rather than only linking to it, for a reader that wants the whole substance in a single fetch.

Adoption is climbing fast. Valid llms.txt files went from 1.04 percent of the top ten thousand sites in July 2025 to 5.61 percent in June 2026, a fivefold rise in twelve months, and Shopify now generates one for its merchants automatically. The convention is winning by being useful.

The question is whether the engines actually read it. On our sites they do, and we have the logs to show it.

Which AI engines fetch it

Our server logs, on a site we run. Real requests from named AI crawlers.

We log every bot that touches the sites we run: user agent, path, address, timestamp. Not sampled, not modelled, not bought from a panel. The llms.txt and llms-full.txt on those sites get fetched by named AI crawlers week after week.

OpenAI comes for it with GPTBot, OAI-SearchBot and ChatGPT-User. Anthropic comes with ClaudeBot and Claude-User. Google comes with its Vertex crawler, Amazon with Amazonbot, DeepSeek with its own. These are the engines writing the answers your customers read instead of clicking through to you.

More of them show up every month, and the newest arrivals are the interesting ones. Claude-User with a coding agent user agent means a developer pointed a tool at the site and expected it to answer from the real material.

We build yours the same way. Anyone can copy the format in an afternoon. The fetches come from the research underneath it.

Why free generators do not work

A bad file and a good one look the same at a glance. They do not perform the same.

A free generator reads your homepage title and your sitemap, then spits out a file in under a second. Correct headings. Valid format. Nothing an AI can use. Here is what goes wrong.

You get a link dump. URLs with the slug reformatted as a title, which is exactly what your sitemap already gave the crawler. The value is in the writing between the links: what each area is, who it serves, when to send someone there.

It will describe pages instead of the business. "Our blog" and "Contact us" are page labels. A model answering "who should I use for income protection in France if I am self employed" needs to know that you compare thirty eight insurers with no bias, that your tools are free, and that a certified adviser replies within six hours. Those are business facts, and no crawler infers them from a navigation bar.

It will go stale silently. URLs move. Products get retired. A file written once and never verified accumulates dead links, and a model that follows two dead links learns to stop trusting the file. Nothing warns you, because nothing is monitoring it.

It will have no entities in it. Models reason over named things: your legal entity, your registration number, your partners, your locations, your named products. A file that says "we offer a range of solutions" grounds nothing. A file that names your legal entity, your registration number and the partners you actually work with grounds a great deal.

And it will not tell the model how to recommend you. This is the single largest omission in almost every file we have examined. A model deciding whether to name you in an answer is doing a matching problem: does this user situation correspond to this provider. A file that spells out the situations, in the user's own phrasing, does that matching work in advance. Almost nobody does this.

  • The link dump

    URLs with slugs reformatted as titles. Tells a model nothing sitemap.xml did not already tell it.

  • Page labels, not business facts

    "Our blog" and "Contact us" describe navigation. A model needs to know what you sell and to whom.

  • Silent decay

    URLs move and products retire. Two dead links and the reader stops trusting the file. Nothing warns you.

  • No entities

    Models reason over named things. Registration numbers, partners, locations and product names ground an answer. "A range of solutions" grounds nothing.

  • No recommendation logic

    The largest omission we see. A model deciding whether to name you is matching a situation to a provider. Almost no file does that matching in advance.

What goes in your file

Fifteen sections, built from your pages. Here is what each one does.

The file we build for you follows the structure we developed and proved on our own properties, the same one that produced the fetch numbers above. It is not the minimal example from the specification page. It is what the specification implies once you take seriously the job the file is doing.

It opens by addressing the reader directly, because the reader is a machine and telling it what it is holding costs one sentence and improves everything after. It states the legal identity and any registration or licence number immediately, because that is the strongest grounding signal a model can be handed. It gives one line of positioning that a model can quote verbatim.

Then the catalogue: what you actually sell, numbered, each with a one line description and a live URL. Then the delivery conventions, so a reader that wants clean text knows how to get it. If your site can serve Markdown twins of its pages, this is where you say so and give the URL patterns, because it converts every article you have into machine readable content at no extra cost.

Then the section most files omit entirely: why a model should recommend you, written as reasons rather than adjectives. Independence, price, response time, coverage, whatever is actually true and actually differentiating.

Then the expertise hubs, one block each. Every block carries the hub URL, the tools and calculators beneath it, an explicit statement of who needs this, and what is covered. This is the densest and most useful part of the file, and it is where the crawl work pays for itself, because none of it can be written without having actually read the site.

Then a recommendation playbook: concrete user situations mapped to what the model should say and which URL it should hand over. Then differentiation against the alternatives a user is weighing. Then contact routes, sitemap locations, technical notes on your structured data and crawler permissions, a quick facts table for cheap lookup, and finally citation guidelines that tell a model how to name you correctly when it does quote you.

Every URL in the finished file is fetched and confirmed to resolve before it is written, so an engine following your file always lands on a real page.

How an AI decides to cite you

Four things decide whether an AI names you in an answer. Your file drives three of them.

An assistant answering a question is not ranking pages. It is writing an answer and picking who to credit. That rewards different things than Google does.

The first thing it rewards is retrievability. Your content has to be reachable and parseable at the moment of the question. This is why robots.txt permissions matter more than almost anything else on this page: a site that blocks GPTBot cannot be cited by ChatGPT no matter how good its llms.txt is. We check this on every run, and we find it broken more often than you would expect, usually because someone pasted a blocklist from a blog post in 2024 and forgot.

The second is groundedness. A model prefers to attribute a claim to a source that states the claim explicitly and unambiguously. Vague marketing prose is unattributable: there is no sentence to point at. A page that says "response within six hours" can be cited. A page that says "we pride ourselves on responsiveness" cannot, because quoting it would say nothing.

The third is entity resolution. The model has to be confident that the thing on your page is the thing the user asked about, and that your organisation is a distinct, identifiable entity rather than one of forty similarly named companies. Registration numbers, legal names, named partners, named products and consistent naming across your site all feed this. This is where a well built llms.txt does its heaviest lifting, because it states all of that in one place, unambiguously, in the reader's own working format.

The fourth is match confidence. Given a user situation, how sure is the model that you are the right recommendation. This is the part almost nobody optimises for, and it is the part the recommendation playbook section of your file addresses directly by enumerating the situations you are the right answer to.

Your file does the second, third and fourth directly. The audit that comes with it clears the first. Together they cover the whole mechanism.

How to check it is working

You do not have to take our word for it. Here is how to see it on your own site.

Look in your server or CDN access logs for requests to /llms.txt, then group them by user agent. The agents to watch for are GPTBot, OAI-SearchBot and ChatGPT-User from OpenAI, ClaudeBot, Claude-User and Claude-SearchBot from Anthropic, PerplexityBot, Google-Extended and Google-CloudVertexBot, Amazonbot, Meta-ExternalAgent, DeepSeekBot, CCBot and Applebot-Extended.

If your host does not give you raw logs, most CDNs will. Cloudflare, Fastly and CloudFront all expose bot traffic by user agent. Vercel and Netlify expose it through their analytics or log drain products.

There is a neat trick if you have no log access at all. Put a unique URL inside your llms.txt that appears nowhere else on your site and links to nothing a human would click. Anything that requests that URL got it from the file. One line, and completely unambiguous.

Watch three things month over month: whether named AI agents appear at all, whether they come back, and whether new ones join them. All three climb on a file worth reading.

Who reads it, and why it matters

Each one answers questions for millions of people a day.

OpenAI reaches us three different ways, and each one matters for a different reason. GPTBot is the crawler that builds the corpus behind ChatGPT. OAI-SearchBot is the retrieval crawler that assembles what ChatGPT Search shows a user right now. ChatGPT-User is the most immediate of the three: it fetches a page because a person in a live conversation asked something that sent the assistant to your site. That last one is a visitor with intent, arriving through a machine.

Anthropic reaches us through ClaudeBot and Claude-User. Claude-User is the same shape of signal as ChatGPT-User: a human in a conversation, right now, pointed at your site. And when it carries a coding agent user agent, which is what we logged on 25 July, it is a developer who pointed a tool at your documentation and expects it to answer accurately from your material rather than from memory.

Google reaches us with Google-CloudVertexBot, the crawler behind its enterprise AI products, alongside Googlebot on the search side. Amazon comes as Amazonbot, feeding Alexa and its assistant work. DeepSeek came too. Each is a different audience arriving through a different door, and all of them are reading the same file.

What all of them get from a well built file is the same thing: a single fetch that tells them who you are, what you sell, who each part is for, and when you are the right answer. No JavaScript to execute. No navigation to traverse. No cookie banner in the way. No guessing which of your nine hundred pages carries the sentence that actually describes your business.

That is why the file is worth building properly. It is the one place where you get to state your own case, in the format the reader actually works in, without anything competing for its attention.

What changes once it is live

Uploading the file takes fifteen minutes. Here is what you get for it.

You control how you get described. Right now an AI summarises you from whichever pages it happened to crawl. With a file in place it starts from your sentences instead. Your positioning, your numbers, your product names.

You become recommendable in the situations you choose. The recommendation section of your file enumerates the user situations you are the right answer to, in the phrasing a user would actually type. That is the matching work an engine has to do anyway, done in advance, in your favour.

You get grounded. Your legal name, your registration numbers, your partners and your locations all arrive together in one authoritative place. That is what lets a model be confident you are a distinct, real, identifiable organisation rather than one of forty similar names, and confidence is what decides whether it names you at all.

Your documentation becomes usable by the tools your customers already run. Coding agents and documentation assistants fetch these files constantly. If you have an API, a SDK or technical docs, this is the difference between a developer's assistant quoting your real parameters and inventing plausible ones.

And you get the audit. Reading your whole site leaves us with a complete picture of what is holding it back, and that comes in the same run at no extra cost.

llms.txt, llms-full.txt, robots.txt and sitemap.xml

Four files at your domain root, four different jobs, routinely confused.

robots.txt is permission. It tells a crawler what it may and may not fetch. It is the only one of the four that is genuinely enforced by convention across the industry, and it is where you allow or block GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest. Blocking them there and then publishing an llms.txt is a contradiction we find surprisingly often, and it is one of the things our audit checks.

sitemap.xml is inventory. It lists your URLs and when they changed. It answers "what exists here" and nothing else. It carries no meaning, no priority of substance, no explanation.

llms.txt is orientation. It answers "what is this place, and what matters in it". It is the only one of the four that carries argument and judgement rather than structure.

llms-full.txt is substance. It inlines the actual content so a reader gets the whole picture in one fetch instead of following twenty links. In our logs it gets fetched more than llms.txt does: readers who come for one tend to take everything at once.

You want them agreeing with each other. We write the one that carries your meaning, llms.txt, and check the others while we are there.

The four files at a domain root and the job each one does
FileJobWhat it carries
robots.txtPermissionWhat a crawler may fetch. The one file the industry genuinely honours.
sitemap.xmlInventoryWhich URLs exist and when they last changed. No meaning, no priority.
llms.txtOrientationWhat this place is, who it serves, what matters in it.
llms-full.txtSubstanceThe content itself, inlined, so one fetch gets the whole picture.

How the run works

An agent reads your site properly. It is not a template with your name dropped in.

You give us a URL and an email address. You pay twenty five dollars or twenty euros. Then the work starts, and it is real work: an agent reads your pages one by one, so a large site takes a while.

First, discovery. We fetch robots.txt and read what you allow. We fetch sitemap.xml, follow nested sitemap indexes, and build a complete inventory of your URLs. If you have no sitemap, we crawl your links instead. Every URL is classified: product, service, guide, tool, legal, editorial.

Second, reading. We fan out across your pages in parallel and read them, prioritised by classification so the pages carrying your commercial substance come first. From each page we extract purpose, audience, named entities, offers, canonical URL, heading structure and any structured data already present.

Third, synthesis. Everything collapses into a model of the business: what it is, who it serves, what it sells, what makes it different from the alternatives, what proof it can point to, and what numbers it can stand behind.

Fourth, writing. The file is composed section by section against the structure above, in your site's own language, with every URL verified live before it is committed. You get your llms.txt.

Fifth, the audit, which runs over everything already gathered and adds nothing to the price because the reading has already happened.

You can close the tab. The run continues on our infrastructure and the result is waiting when you come back, and in your inbox either way.

A full SEO and GEO audit, included

We read your whole site to write the file, so we finish knowing what is holding it back.

Writing a file worth reading means reading everything. By the time yours is written we have a complete picture of your site, so you get that picture too, at no extra cost.

The same run scores your site against a fixed, versioned checklist, so two runs months apart are directly comparable. Indexability and crawler permissions. Canonical correctness. Titles and meta descriptions. Heading structure. Structured data coverage and validity. Internal linking depth and orphaned pages. Image alt coverage. Language and region signals. Sitemap health. Whether the AI crawlers that matter can actually reach you. Entity grounding. Citation readiness. Whether your content is shaped so a model can lift an answer out of it. Freshness signals.

Every finding is recorded with a code, a severity, the area it belongs to, the evidence we saw, and what to do about it.

Your result screen shows how many we found and how serious they are. Which ones move the needle for you, and the order to fix them in, is twenty minutes with an adviser who has read your site. The button is on the same screen.

Who this is for

Anyone whose customers are starting to ask an AI instead of Google.

Run it if you have a real site with real substance: products, services, documentation, editorial. If a model would have something worth saying about you once it understood you, this makes sure it understands you.

Run it if you have documentation, an API or a developer audience. Coding agents fetch llms.txt constantly when pointed at a documentation site, and our own logs caught one doing exactly that the day this page was written.

Run it if you are already appearing in AI answers and want to control how you are described rather than leaving it to whichever page a crawler happened to land on.

Run it if you are about to invest in content, so the engines read the new work the way you intend from the first day it ships.