Skip to main content Scroll Top

llms.txt Generator: Build the File, and Know What It Is Worth

Updated September 2026 · Written and maintained by the Progression Agency strategy team

Fill in a short form and get a valid llms.txt to copy or download. Then read the honest assessment of what publishing one actually does.

On this page · 10 sections
  1. What this file is, stated plainly
  2. The part that pays now
  3. How the generator handles the details
  4. Where to put it, and what to serve it as
  5. Keeping it honest over time
  6. Who this file is actually for
  7. What we would not do with this file
  8. The reference tables, in one place
  9. Everything else we have written on search, AI and getting found
  10. Validating an llms.txt file, and what “valid” even means here

The short answerllms.txt is a proposed convention, not a standard: no specification body, no search engine documenting that it reads the file, no assistant committing to consume it. It is still worth twenty minutes, because producing it forces you to decide which fifteen pages are your best and to describe each in one line — and those descriptions belong on the pages themselves, where they do work today. Publish the file because it is free, not because it is proven.

Nothing here claims the file is read by any specific system. If that changes we will update this page rather than leave the old framing standing.

llms.txt generator

Fill in the fields and get a valid, well-formed llms.txt you can copy or download. Read the section underneath it before you decide whether to publish one.

Results appear here.

Runs entirely in your browser. Nothing you paste is uploaded, logged or stored — which is also why it takes pasted text rather than a URL: a browser cannot read another site’s files without that site’s permission.

What this file is, stated plainly

A proposed convention with no specification body, no search engine documenting that it reads the file, and no assistant committing to consume it.

It is a proposal, not a standard

llms.txt is a suggested convention: a markdown file at the root of a site listing the pages its owner considers most worth reading, with a short description of each. It has no specification body behind it, no search engine documents consuming it, and no assistant publicly commits to reading it. Anyone telling you it is a requirement is selling something.

Which does not make it worthless

It costs minutes, it carries no risk, and if a consumer emerges you already have one. More usefully, producing it forces a decision most organisations have never made explicitly: which fifteen pages would you send a new customer to, and what would you say about each. That exercise pays regardless of whether the file is ever read.

We would rather say this than sell it

There is a version of this page that implies publishing a text file will get you cited. It would convert better and it would be untrue. What we will say is that the curation exercise is worth doing, the file is free, and the descriptions you write belong on the pages themselves where they do work today.

What publishing an llms.txt does and does not doWhat publishing an llms.txt does and does not do
Three yeses and five noes. We build the generator because the curation exercise is genuinely useful and the file costs nothing; we would not sell it as a visibility lever.

The part that pays now

The one-line descriptions you write for this file belong on the pages themselves, where they do work today whether or not the file is ever read.

The descriptions belong on the pages

A one-line description of what a page answers is exactly the kind of text that helps a page be extracted and summarised correctly — as an opening sentence, as a meta description, as the answer paragraph directly under the heading. Most sites have never written one. Writing fifteen of them for this file and then leaving them only in the file wastes the work.

The pages you cannot describe are the finding

In our experience the exercise stalls on two or three pages every time, and they are always the same kind: pages that exist because a competitor had one, pages assembled from a brief nobody remembers, pages that are three topics wearing one heading. Being unable to write a single line about a page is a content judgement, delivered free.

Fifteen is usually the right size

This is a shortlist. An XML sitemap already handles the complete inventory and is a documented mechanism with known consumers. A curated file that lists everything is just a worse sitemap, and it loses the only property that makes it interesting.

Proposal — Not a standard. No specification body.
Uncertain — Consumer set. Nobody documents reading it.
Cheap — To publish. Minutes, no risk.
Useful — As an exercise. Curation is the value.
Curated — Not complete. A sitemap does inventory.
Current — Or worthless. A stale list misleads.

How the generator handles the details

Relative paths, pipe-separated entries and deliberate warnings rather than silent corrections — plus nothing leaving your browser.

Relative paths are expanded if you give a site URL

Enter your domain and any path written as /costs becomes a full URL. Absolute URLs are safer in a file like this because there is no documented base-URL behaviour to rely on, and a consumer that resolves them differently would resolve them wrongly.

Each line takes a title, a path and an optional description

Separated by pipes, which is the least ambiguous separator for text people paste from spreadsheets. A line with only a title still produces a valid entry; a line with a description produces the form the convention actually suggests.

It warns rather than silently fixing

No site URL, no links at all, or more than sixty links each produce a warning instead of a quiet adjustment. A tool that corrects your input without telling you teaches you nothing about your own file.

Everything stays in the browser

There is no upload and no request. You can generate, copy or download, and nothing about your site reaches us.

Why we still think it is worth twenty minutesWhy we still think it is worth twenty minutes
Every client we have run this exercise with found at least one page they could not describe in a sentence. That is the finding, not the file.

Where to put it, and what to serve it as

At the root of the domain, as plain text, returning a 200 — and worth loading yourself afterwards to confirm all three.

At the root, as plain text

The convention places it at the root of the domain — the same level as robots.txt — served with a plain text content type. A file served as HTML, or behind a redirect, or returning anything other than a 200, is not doing its job even if a consumer does eventually look.

Check it the way you would check robots.txt

Load the URL in a browser and confirm you see plain text, not a themed page and not a download prompt. Several hosting setups will happily serve an unknown extension as an attachment, which is technically a file at the right URL and practically useless.

Keep it out of your sitemap and out of your navigation

It is a machine-facing file, not a page. Linking to it from navigation invites confused human visitors and adds nothing for anyone else.

Name — Your organisation. Plainly.
Summary — One sentence. No adjectives.
Core — Your best pages. Five to fifteen.
Docs — Explainers. Where they exist.
Policy — Terms and rules. Short list.
Contact — One line. Real address.

Keeping it honest over time

The entire premise is that a human chose these pages deliberately, which a list pointing at two 404s actively contradicts.

A stale curated list is worse than no list

The entire premise of the file is that a human chose these pages deliberately. A list still pointing at last year’s offers, a retired service and two 404s says the opposite, and it says it in a file specifically framed as your considered recommendation.

Review quarterly, and after any restructure

URL changes are the usual killer. Any migration, any slug change and any content consolidation should include a pass over this file, exactly as it should include a pass over your internal links and your sitemap.

Prune before you add

The temptation with a curated list is to keep appending. If a page has been added and nothing removed for a year, it has quietly become an inventory. Reviewing what to cut is the harder and more useful half of the exercise.

Where this fits with everything else

It sits at the cheap, unproven end. The moves that are actually evidenced and what an audit covers are where the effort belongs first.

Where the effort actually paysWhere the effort actually pays
The top three are worth doing whether or not you ever publish the file. That ordering is the honest version of this topic.

Who this file is actually for

Not your visitors and not search engines, which already have a documented mechanism for this job.

Not for your visitors

Nobody browsing your site will ever open it, and nothing about it improves the experience of a human reader. Treating it as a marketing asset — announcing it, linking it, putting it in a footer — misunderstands what it is and adds clutter for people who do not want it.

Not for search engines either

Search engines already have a mechanism for this and it is the XML sitemap, which they document, discover automatically and act on. This file adds nothing to that relationship and does not replace any part of it.

For a category of consumer that may or may not arrive

The honest framing is that it is a bet with a very low stake. If a major assistant starts reading these files, you already have a good one. If none ever does, you spent twenty minutes and got a curation exercise out of it. That is the whole case, and it is enough.

Sitemap — Documented. Known consumers.
robots.txt — Documented. Known consumers.
Schema — Documented. Known consumers.
llms.txt — Proposed. Unknown consumers.
Pages — The real surface. Where answers come from.
Entity records — Off your site. Also read.

Want this done for your site?We build and maintain the search, content and paid programmes described on this page.

Get a free proposal

What we would not do with this file

Three things commonly sold around llms.txt that we would decline, and the reasoning behind each.

We would not charge for it

It is a text file produced from a form. Any agency billing a setup fee for one is charging for the twenty minutes of thinking, which you can do better than they can because it is your business.

We would not promise anything from publishing it

There is no measured outcome to point at, and inventing one would be exactly the behaviour we criticise elsewhere on this site. If a client asks whether it will help, the answer is that it cannot hurt and nobody can currently show that it helps.

We would not let it replace the work that is evidenced

Crawler access, extractable content, entity consistency and answer-first structure all have observable effects. This file does not. If effort is limited, it should go to the first four and this one should wait for a quiet afternoon.

llms.txt against the files that already have consumersllms.txt against the files that already have consumers
This is the comparison most llms.txt advice leaves out. A sitemap is a documented mechanism with known consumers; llms.txt is a proposal with an uncertain audience.

The reference tables, in one place

Four tables follow: the file’s structure, how it compares with the files that already have consumers, what the generator warns about, and the review triggers. They summarise the sections above and are meant to be usable on their own.

What goes in each part of the file
PartContentsWhy it is thereCommon mistake
Title lineYour organisation nameIdentifies whose file this isA tagline instead of a name
SummaryOne sentence on what you doThe only description a reader may seeMarketing adjectives
NotesDates, caveats, review statusMakes the file checkableLeaving it undated
Core pagesYour five to fifteen bestThe point of the fileListing everything
DocumentationExplainers and guidesDeeper readingDuplicating core pages
PoliciesTerms, privacy, cancellationFrequently asked, rarely linkedOmitting them
ContactOne real addressSomewhere to go nextA contact form URL with no address
llms.txt against the files that already have consumers
FileStatusWho documents consuming itWhat it is for
XML sitemapEstablishedMajor search enginesComplete URL inventory
robots.txtEstablishedMajor search engines and AI operatorsCrawl permissions
Structured dataEstablishedMajor search enginesMachine-readable facts about a page
llms.txtProposedNobody publiclyA curated reading list
Your pagesEstablishedEverythingWhere answers actually come from
Directory recordsEstablishedSearch and assistantsFacts about you off your own site
Press and reviewsEstablishedSearch and assistantsThird-party evidence
What the generator warns about
WarningTriggerWhy it matters
No site URLSite field left blankRelative paths stay relative, with no defined base
No linksNothing in any sectionA file with no links is a header, not an index
Too many linksMore than sixtyCuration is the only property that distinguishes it from a sitemap
Undescribed entriesA line with no descriptionThe description is where the value is
Missing summarySummary field blankThe one line a consumer is most likely to use
No contactContact field blankCheap to add, frequently wanted
No notesNotes field blankAn undated file cannot be judged for freshness
When to review the file
TriggerWhyAction
Any URL or slug changeListed URLs break silentlyRegenerate
Any content consolidationMerged pages leave dead entriesRegenerate
A new flagship pageThe list is a recommendationAdd, and cut something
A retired serviceA stale recommendation misleadsRemove
QuarterlyDrift is invisible otherwiseReview and re-date
After a migrationRoot files are often lost entirelyConfirm it still serves
If a consumer is ever announcedThe file becomes load-bearingReview properly then
Where llms.txt sits among AI-visibility workWhere llms.txt sits among AI-visibility work
Bottom-left is cheap and unproven, which is exactly where this file belongs. It is worth doing because it is cheap, not because it is proven.
A sensible way to use the generatorA sensible way to use the generator
Step five is the one that pays now. Everything before it is preparation for a file that may or may not ever be read.
The honest numbersThe honest numbers
We would rather publish the first number on our own tool page than let a client discover it later.
Sitemap replacement — No. Different jobs.
Ranking lever — No. No evidence.
Set and forget — No. Staleness misleads.
Complete inventory — No. Curation is the point.
Announcement — Pointless. Nobody is waiting.
Paid service — Unjustifiable. It is a text file.
Pick — Your best pages. The hard part.
Describe — One line each. Exposes weak pages.
Summarise — The organisation. One sentence.
Generate — And download. Seconds.
Upload — To your web root. Served as text.
Reuse — Descriptions on pages. Where they pay now.
Honest — About evidence. No guarantees.
Short — A shortlist. Not an inventory.
Absolute — URLs. Not relative paths.
Dated — Reviewed quarterly. Staleness misleads.
Reused — On the pages. The real return.
Free — Always. It is a text file.

Everything else we have written on search, AI and getting found

Validating an llms.txt file, and what “valid” even means here

There is no specification body for llms.txt, which means there is no formal validator and cannot be one in the sense that HTML or JSON Schema have validators.

What can be checked is whether the file is well-formed Markdown, whether its links resolve, whether the structure matches the convention as proposed, and whether it contradicts your own site. That is what this tool checks, and it is worth being precise that this is convention-conformance rather than validation against a standard.

What is genuinely checkable

  • Well-formed Markdown, since the proposal is a Markdown document.
  • An H1 carrying the site or project name.
  • A blockquote summary directly beneath it.
  • Link sections as H2 headings with Markdown link lists underneath.
  • Links that actually resolve rather than pointing at moved or missing pages.
  • Descriptions that say something specific rather than repeating the page title.

What no validator can tell you

  • Whether any AI system reads your file, because none has committed to doing so.
  • Whether it affects how you are described, because nothing has demonstrated that it does.
  • Whether it will matter later, which is unknowable and is the honest reason to keep it cheap.

Does llms.txt actually work?

No search engine or assistant operator has documented reading it. That is the state of it, and anyone telling you otherwise is describing a hope. The reasonable position is that it costs almost nothing to publish a correct one, so the expected value is positive even though the current measured effect is zero. Publish it, keep it accurate, and do not reorganise your content strategy around it.

llms.txt against the files it gets compared to
llms.txtrobots.txtsitemap.xml
Has a specification bodyNoYes, a long-standing conventionYes, sitemaps.org
Documented as read by GoogleNoYesYes
Documented as read by AI assistantsNoYes, for crawler accessNot stated
Controls anythingNoCrawler accessDiscovery
Cost to publishMinutesMinutesUsually automatic
Reasonable to publish anywayYesRequiredYes

AI, AEO and what is changing

Websites and design

Choosing and working with an agency

Social, content and brand

By industry and by situation

Want the pages behind the list looked at

The list is only as good as what it points to. We audit the pages themselves — extractability, answer structure, entity consistency and the access that has to work first.

/contact

Not sure which of these applies to you?Tell us the situation and we will say plainly what we would do first, and what we would not.

Talk it through

AI, AEO and what is changing

Frequently asked questions

Is there an llms.txt validator?
Not in the formal sense, because there is no specification body to validate against. What can be checked is whether the file is well-formed Markdown, follows the proposed structure, and contains links that actually resolve. This tool checks those.
Does llms.txt actually work?
No search engine or assistant operator has documented reading it. The measured effect today is zero. It costs minutes to publish a correct one, which is the only honest argument for doing so.
Is llms.txt a ranking signal?
No. Nothing has been documented or demonstrated to that effect, and any tool or agency claiming otherwise is describing a hope rather than a finding.
What is the difference between llms.txt and robots.txt?
robots.txt is a long-standing convention that controls crawler access and is documented as read by Google and the major AI crawlers. llms.txt is a proposal with no specification body that controls nothing and has not been documented as read by anyone.
How do I check my llms.txt is correct?
Confirm it is well-formed Markdown with an H1 naming the site, a blockquote summary, H2 sections and Markdown link lists, and that every link resolves. That is conformance to the convention as proposed, which is as far as checking can honestly go.
Is llms.txt an official standard?
No. It is a proposed convention with no specification body behind it. No search engine documents consuming it and no assistant publicly commits to reading it.
Then why publish one?
Because it costs minutes, carries no risk, and the exercise of producing it forces a decision most organisations have never made explicitly: which pages are your best, and what would you say about each.
Will it improve my rankings?
There is no evidence that it does, and we would not tell you otherwise. If a page ranks better after you publish one, the cause is almost certainly something else you changed.
Does it replace my sitemap?
No, and it should not try. A sitemap is a documented mechanism with known consumers that handles complete URL inventory. This file is a curated shortlist, and a curated file listing everything is just a worse sitemap.
How many links should it have?
Usually five to fifteen. Past about sixty you have lost the only property that makes the file interesting, which is that a human chose these deliberately.
Where does the file go?
At the root of your domain, the same level as robots.txt, served as plain text with a 200 response. Load the URL yourself afterwards and confirm you see text rather than a themed page or a download prompt.
Should I link to it from my navigation?
No. It is a machine-facing file, not a page. Linking it invites confused human visitors and adds nothing for anyone else.
What is the most useful part of doing this?
Writing the one-line descriptions, then putting them on the pages themselves. That is where they do work today — as opening sentences, meta descriptions and answer paragraphs under headings.
What if I cannot describe one of my pages?
That is a content finding, delivered free. In our experience the exercise stalls on the same kinds of pages every time: ones that exist because a competitor had one, or that are three topics under a single heading.
Should the URLs be absolute or relative?
Absolute. There is no documented base-URL behaviour to rely on, so a consumer resolving relative paths differently would resolve them wrongly. The generator expands relative paths for you if you supply a site URL.
How do I format each line?
Title, then a pipe, then the path or URL, then optionally another pipe and a description. A line with only a title still produces a valid entry.
Does the tool upload anything?
No. It runs entirely in your browser, there is no request to any server, and you can confirm that in your browser’s network panel while using it.
Can I download the result?
Yes, as a file named llms.txt, or copy it to your clipboard. Both happen locally.
How often should I update it?
Quarterly, and immediately after any URL change, content consolidation or migration. A stale curated list is worse than none, because the file is explicitly framed as your considered recommendation.
What is the most common way it breaks?
URL changes. The file sits outside the CMS, so nothing updates it when slugs change, and nothing alerts you when its links start 404ing.
Should I include policies?
Yes. Cancellation terms, privacy and similar pages are frequently asked about and rarely linked prominently, so a curated list is a reasonable place to surface them.
Should I announce that I published it?
There is nobody waiting for the announcement. Publish it, note the date, move on.
Is there any risk in publishing one?
None that we can identify. It exposes no information you have not already published and it cannot affect crawling or indexing.
Would you charge for setting one up?
No, and we would question anyone who does. It is a text file, and this generator produces it in about twenty minutes of your own thinking time.
Does it help with AI Overviews?
There is no evidence that it does. Those are built from the search index, which is governed by crawling, indexing and content quality rather than by a root text file.
What should I do instead if I only have an hour?
Check crawler access, then make your highest-intent page answer its question in the first two sentences. Both are evidenced; this file is not.
Will you update this page if that changes?
Yes. If a major consumer documents reading the file, that changes the honest answer and we will say so here rather than quietly leaving the old framing up.

Want this done for your site?We build and maintain the search, content and paid programmes described on this page.

Get a free proposal

Get a free marketing proposal

Tell us what you are trying to grow and we will come back with a plan, not a pitch deck. Same-day reply on weekdays.

Privacy Preferences
When you visit our website, it may store information through your browser from specific services, usually in form of cookies. Here you can change your privacy preferences. Please note that blocking some types of cookies may impact your experience on our website and the services we offer.
Contact Us