Skip to main content Scroll Top

llms.txt Generator: Build the File, and Know What It Is Worth

Updated September 2026 · Written and maintained by the Progression Agency strategy team

Fill in a short form and get a valid llms.txt to copy or download. Then read the honest assessment of what publishing one actually does.

On this page · 10 sections
  1. What this file is, stated plainly
  2. The part that pays now
  3. How the generator handles the details
  4. Where to put it, and what to serve it as
  5. Keeping it honest over time
  6. Who this file is actually for
  7. What we would not do with this file
  8. The reference tables, in one place
  9. Everything else we have written on search, AI and getting found
  10. Video: site files, crawling and AI access

The short answerllms.txt is a proposed convention, not a standard: no specification body, no search engine documenting that it reads the file, no assistant committing to consume it. It is still worth twenty minutes, because producing it forces you to decide which fifteen pages are your best and to describe each in one line — and those descriptions belong on the pages themselves, where they do work today. Publish the file because it is free, not because it is proven.

Nothing here claims the file is read by any specific system. If that changes we will update this page rather than leave the old framing standing.

llms.txt generator

Fill in the fields and get a valid, well-formed llms.txt you can copy or download. Read the section underneath it before you decide whether to publish one.

Results appear here.

Runs entirely in your browser. Nothing you paste is uploaded, logged or stored, and there is no network request — which is also why it takes pasted text rather than a URL: a browser cannot read another site’s files without that site’s permission.

What this file is, stated plainly

A proposed convention with no specification body, no search engine documenting that it reads the file, and no assistant committing to consume it.

It is a proposal, not a standard

llms.txt is a suggested convention: a markdown file at the root of a site listing the pages its owner considers most worth reading, with a short description of each. It has no specification body behind it, no search engine documents consuming it, and no assistant publicly commits to reading it. Anyone telling you it is a requirement is selling something.

Which does not make it worthless

It costs minutes, it carries no risk, and if a consumer emerges you already have one. More usefully, producing it forces a decision most organisations have never made explicitly: which fifteen pages would you send a new customer to, and what would you say about each. That exercise pays regardless of whether the file is ever read.

We would rather say this than sell it

There is a version of this page that implies publishing a text file will get you cited. It would convert better and it would be untrue. What we will say is that the curation exercise is worth doing, the file is free, and the descriptions you write belong on the pages themselves where they do work today.

The part that pays now

The one-line descriptions you write for this file belong on the pages themselves, where they do work today whether or not the file is ever read.

The descriptions belong on the pages

A one-line description of what a page answers is exactly the kind of text that helps a page be extracted and summarised correctly — as an opening sentence, as a meta description, as the answer paragraph directly under the heading. Most sites have never written one. Writing fifteen of them for this file and then leaving them only in the file wastes the work.

The pages you cannot describe are the finding

In our experience the exercise stalls on two or three pages every time, and they are always the same kind: pages that exist because a competitor had one, pages assembled from a brief nobody remembers, pages that are three topics wearing one heading. Being unable to write a single line about a page is a content judgement, delivered free.

Fifteen is usually the right size

This is a shortlist. An XML sitemap already handles the complete inventory and is a documented mechanism with known consumers. A curated file that lists everything is just a worse sitemap, and it loses the only property that makes it interesting.

An XML sitemap already handles the complete inventory and is a documented mechanism with known consumers.

How the generator handles the details

Relative paths, pipe-separated entries and deliberate warnings rather than silent corrections — plus nothing leaving your browser.

Relative paths are expanded if you give a site URL

Enter your domain and any path written as /costs becomes a full URL. Absolute URLs are safer in a file like this because there is no documented base-URL behaviour to rely on, and a consumer that resolves them differently would resolve them wrongly.

Each line takes a title, a path and an optional description

Separated by pipes, which is the least ambiguous separator for text people paste from spreadsheets. A line with only a title still produces a valid entry; a line with a description produces the form the convention actually suggests.

It warns rather than silently fixing

No site URL, no links at all, or more than sixty links each produce a warning instead of a quiet adjustment. A tool that corrects your input without telling you teaches you nothing about your own file.

Everything stays in the browser

There is no upload and no request. You can generate, copy or download, and nothing about your site reaches us.

Where to put it, and what to serve it as

At the root of the domain, as plain text, returning a 200 — and worth loading yourself afterwards to confirm all three.

At the root, as plain text

The convention places it at the root of the domain — the same level as robots.txt — served with a plain text content type. A file served as HTML, or behind a redirect, or returning anything other than a 200, is not doing its job even if a consumer does eventually look.

Check it the way you would check robots.txt

Load the URL in a browser and confirm you see plain text, not a themed page and not a download prompt. Several hosting setups will happily serve an unknown extension as an attachment, which is technically a file at the right URL and practically useless.

Keep it out of your sitemap and out of your navigation

It is a machine-facing file, not a page. Linking to it from navigation invites confused human visitors and adds nothing for anyone else.

Linking to it from navigation invites confused human visitors and adds nothing for anyone else.

Keeping it honest over time

The entire premise is that a human chose these pages deliberately, which a list pointing at two 404s actively contradicts.

A stale curated list is worse than no list

The entire premise of the file is that a human chose these pages deliberately. A list still pointing at last year’s offers, a retired service and two 404s says the opposite, and it says it in a file specifically framed as your considered recommendation.

Review quarterly, and after any restructure

URL changes are the usual killer. Any migration, any slug change and any content consolidation should include a pass over this file, exactly as it should include a pass over your internal links and your sitemap.

Prune before you add

The temptation with a curated list is to keep appending. If a page has been added and nothing removed for a year, it has quietly become an inventory. Reviewing what to cut is the harder and more useful half of the exercise.

Where this fits with everything else

It sits at the cheap, unproven end. The moves that are actually evidenced and what an audit covers are where the effort belongs first.

Who this file is actually for

Not your visitors and not search engines, which already have a documented mechanism for this job.

Not for your visitors

Nobody browsing your site will ever open it, and nothing about it improves the experience of a human reader. Treating it as a marketing asset — announcing it, linking it, putting it in a footer — misunderstands what it is and adds clutter for people who do not want it.

Not for search engines either

Search engines already have a mechanism for this and it is the XML sitemap, which they document, discover automatically and act on. This file adds nothing to that relationship and does not replace any part of it.

For a category of consumer that may or may not arrive

The honest framing is that it is a bet with a very low stake. If a major assistant starts reading these files, you already have a good one. If none ever does, you spent twenty minutes and got a curation exercise out of it. That is the whole case, and it is enough.

Search engines already have a mechanism for this and it is the XML sitemap, which they document, discover automatically and act on.

What we would not do with this file

Three things commonly sold around llms.txt that we would decline, and the reasoning behind each.

We would not charge for it

It is a text file produced from a form. Any agency billing a setup fee for one is charging for the twenty minutes of thinking, which you can do better than they can because it is your business.

We would not promise anything from publishing it

There is no measured outcome to point at, and inventing one would be exactly the behaviour we criticise elsewhere on this site. If a client asks whether it will help, the answer is that it cannot hurt and nobody can currently show that it helps.

We would not let it replace the work that is evidenced

Crawler access, extractable content, entity consistency and answer-first structure all have observable effects. This file does not. If effort is limited, it should go to the first four and this one should wait for a quiet afternoon.

Want this done for your site?We build and maintain the search, content and paid programmes described on this page.

Get a free proposal

The reference tables, in one place

Four tables follow: the file’s structure, how it compares with the files that already have consumers, what the generator warns about, and the review triggers. They summarise the sections above and are meant to be usable on their own.

What goes in each part of the file
PartContentsWhy it is thereCommon mistake
Title lineYour organisation nameIdentifies whose file this isA tagline instead of a name
SummaryOne sentence on what you doThe only description a reader may seeMarketing adjectives
NotesDates, caveats, review statusMakes the file checkableLeaving it undated
Core pagesYour five to fifteen bestThe point of the fileListing everything
DocumentationExplainers and guidesDeeper readingDuplicating core pages
PoliciesTerms, privacy, cancellationFrequently asked, rarely linkedOmitting them
ContactOne real addressSomewhere to go nextA contact form URL with no address
llms.txt against the files that already have consumers
FileStatusWho documents consuming itWhat it is for
XML sitemapEstablishedMajor search enginesComplete URL inventory
robots.txtEstablishedMajor search engines and AI operatorsCrawl permissions
Structured dataEstablishedMajor search enginesMachine-readable facts about a page
llms.txtProposedNobody publiclyA curated reading list
Your pagesEstablishedEverythingWhere answers actually come from
Directory recordsEstablishedSearch and assistantsFacts about you off your own site
Press and reviewsEstablishedSearch and assistantsThird-party evidence
What the generator warns about
WarningTriggerWhy it matters
No site URLSite field left blankRelative paths stay relative, with no defined base
No linksNothing in any sectionA file with no links is a header, not an index
Too many linksMore than sixtyCuration is the only property that distinguishes it from a sitemap
Undescribed entriesA line with no descriptionThe description is where the value is
Missing summarySummary field blankThe one line a consumer is most likely to use
No contactContact field blankCheap to add, frequently wanted
No notesNotes field blankAn undated file cannot be judged for freshness
When to review the file
TriggerWhyAction
Any URL or slug changeListed URLs break silentlyRegenerate
Any content consolidationMerged pages leave dead entriesRegenerate
A new flagship pageThe list is a recommendationAdd, and cut something
A retired serviceA stale recommendation misleadsRemove
QuarterlyDrift is invisible otherwiseReview and re-date
After a migrationRoot files are often lost entirelyConfirm it still serves
If a consumer is ever announcedThe file becomes load-bearingReview properly then

Everything else we have written on search, AI and getting found

AI, AEO and what is changing

Websites and design

Choosing and working with an agency

Social, content and brand

By industry and by situation

Want the pages behind the list looked at

The list is only as good as what it points to. We audit the pages themselves — extractability, answer structure, entity consistency and the access that has to work first.

/contact

Not sure which of these applies to you?Tell us the situation and we will say plainly what we would do first, and what we would not.

Talk it through

Video: site files, crawling and AI access

Background viewing only. The assessment above is ours and is written out in full.

AI, AEO and what is changing

What publishing an llms.txt does and does not do
Three yeses and five noes. We build the generator because the curation exercise is genuinely useful and the file costs nothing; we would not sell it as a visibility lever.
Why we still think it is worth twenty minutes
Every client we have run this exercise with found at least one page they could not describe in a sentence. That is the finding, not the file.
Where the effort actually pays
The top three are worth doing whether or not you ever publish the file. That ordering is the honest version of this topic.
llms.txt against the files that already have consumers
This is the comparison most llms.txt advice leaves out. A sitemap is a documented mechanism with known consumers; llms.txt is a proposal with an uncertain audience.
Where llms.txt sits among AI-visibility work
Bottom-left is cheap and unproven, which is exactly where this file belongs. It is worth doing because it is cheap, not because it is proven.
A sensible way to use the generator
Step five is the one that pays now. Everything before it is preparation for a file that may or may not ever be read.
The honest numbers
We would rather publish the first number on our own tool page than let a client discover it later.

Frequently asked questions

Is llms.txt an official standard?
No. It is a proposed convention with no specification body behind it. No search engine documents consuming it and no assistant publicly commits to reading it.
Then why publish one?
Because it costs minutes, carries no risk, and the exercise of producing it forces a decision most organisations have never made explicitly: which pages are your best, and what would you say about each.
Will it improve my rankings?
There is no evidence that it does, and we would not tell you otherwise. If a page ranks better after you publish one, the cause is almost certainly something else you changed.
Does it replace my sitemap?
No, and it should not try. A sitemap is a documented mechanism with known consumers that handles complete URL inventory. This file is a curated shortlist, and a curated file listing everything is just a worse sitemap.
How many links should it have?
Usually five to fifteen. Past about sixty you have lost the only property that makes the file interesting, which is that a human chose these deliberately.
Where does the file go?
At the root of your domain, the same level as robots.txt, served as plain text with a 200 response. Load the URL yourself afterwards and confirm you see text rather than a themed page or a download prompt.
Should I link to it from my navigation?
No. It is a machine-facing file, not a page. Linking it invites confused human visitors and adds nothing for anyone else.
What is the most useful part of doing this?
Writing the one-line descriptions, then putting them on the pages themselves. That is where they do work today — as opening sentences, meta descriptions and answer paragraphs under headings.
What if I cannot describe one of my pages?
That is a content finding, delivered free. In our experience the exercise stalls on the same kinds of pages every time: ones that exist because a competitor had one, or that are three topics under a single heading.
Should the URLs be absolute or relative?
Absolute. There is no documented base-URL behaviour to rely on, so a consumer resolving relative paths differently would resolve them wrongly. The generator expands relative paths for you if you supply a site URL.
How do I format each line?
Title, then a pipe, then the path or URL, then optionally another pipe and a description. A line with only a title still produces a valid entry.
Does the tool upload anything?
No. It runs entirely in your browser, there is no request to any server, and you can confirm that in your browser’s network panel while using it.
Can I download the result?
Yes, as a file named llms.txt, or copy it to your clipboard. Both happen locally.
How often should I update it?
Quarterly, and immediately after any URL change, content consolidation or migration. A stale curated list is worse than none, because the file is explicitly framed as your considered recommendation.
What is the most common way it breaks?
URL changes. The file sits outside the CMS, so nothing updates it when slugs change, and nothing alerts you when its links start 404ing.
Should I include policies?
Yes. Cancellation terms, privacy and similar pages are frequently asked about and rarely linked prominently, so a curated list is a reasonable place to surface them.
Should I announce that I published it?
There is nobody waiting for the announcement. Publish it, note the date, move on.
Is there any risk in publishing one?
None that we can identify. It exposes no information you have not already published and it cannot affect crawling or indexing.
Would you charge for setting one up?
No, and we would question anyone who does. It is a text file, and this generator produces it in about twenty minutes of your own thinking time.
Does it help with AI Overviews?
There is no evidence that it does. Those are built from the search index, which is governed by crawling, indexing and content quality rather than by a root text file.
What should I do instead if I only have an hour?
Check crawler access, then make your highest-intent page answer its question in the first two sentences. Both are evidenced; this file is not.
Will you update this page if that changes?
Yes. If a major consumer documents reading the file, that changes the honest answer and we will say so here rather than quietly leaving the old framing up.

Sources and further reading

  1. Google Search Essentials — SEO starter guide
  2. Google: creating helpful, reliable, people-first content
  3. Google: intro to structured data
  4. Google: LocalBusiness structured data
  5. Google: FAQPage structured data
  6. Google: Article structured data
  7. Google: Product structured data
  8. Google: title links in search results
  9. Google: control your snippets
  10. Google: robots.txt introduction
  11. Google: sitemaps overview
  12. Google: consolidate duplicate URLs
  13. Google: redirects and Search
  14. Google: JavaScript SEO basics
  15. Google: multi-regional and multilingual sites
  16. Google Search Central Blog
  17. Google: get started with Search Console
  18. Google: how local search results are determined
  19. Google Business Profile: prohibited and restricted content
  20. Google Business Profile: address and service area guidelines
  21. Google Business Profile: review policy
  22. Google Business Profile: add or edit categories
  23. Google Ads: location targeting settings
  24. Google Ads: about negative keywords
  25. Google Ads: about Quality Score
  26. Google Ads: importing offline conversions
  27. Google Ads: about Smart Bidding
  28. Google Ads: about Performance Max
  29. Google Local Services Ads: eligibility and screening
  30. Google Ads: keyword match types
  31. Google Analytics 4: about conversions
  32. Google Analytics 4: attribution models
  33. US Census Bureau QuickFacts: New Jersey
  34. US Census Bureau: American Community Survey
  35. US Census: Statistics of US Businesses
  36. Bureau of Labor Statistics: New Jersey data
  37. BLS: Occupational Employment and Wage Statistics
  38. NJ Department of Labor: labor market information
  39. New Jersey Business Action Center
  40. US Small Business Administration: New Jersey district
  41. USA.gov: business resources
  42. web.dev: Core Web Vitals explained
  43. web.dev: Largest Contentful Paint
  44. web.dev: Cumulative Layout Shift
  45. web.dev: Interaction to Next Paint
  46. Google PageSpeed Insights
  47. Google Rich Results Test
  48. Google Search Console
  49. W3C Markup Validation Service
  50. Schema.org: LocalBusiness type
  51. Schema.org: Service type
  52. Schema.org: FAQPage type
  53. Schema.org: HowTo type
  54. W3C: WCAG 2.2 quick reference
  55. FTC: CAN-SPAM Act compliance guide
  56. FCC: telemarketing and robocall rules (TCPA)
  57. FTC endorsement guides — reviews and testimonials
  58. FTC: rule on consumer reviews and testimonials
  59. HHS: HIPAA guidance on online tracking technologies
  60. New Jersey Courts: attorney advertising guidelines
  61. New Jersey DCA: construction codes and permits
  62. New Jersey Home Improvement Contractor registration
  63. New Jersey Division of Consumer Affairs
  64. TikTok for Business
  65. TikTok Creative Center
  66. TikTok Ads Help Center
  67. TikTok Community Guidelines
  68. TikTok Terms of Service
  69. TikTok Privacy Policy
  70. TikTok Safety Center
  71. TikTok Transparency Center
  72. TikTok Creator Portal
  73. TikTok Newsroom
  74. TikTok for Developers
  75. TikTok advertising solutions
  76. TikTok Creator Marketplace
  77. TikTok Business Center
  78. TikTok for Business blog
  79. TikTok Creative Center: top ads
  80. TikTok Branded Content policy
  81. TikTok Shop for sellers
  82. Instagram for Business
  83. Instagram for Creators
  84. Instagram Help Center
  85. About Instagram
  86. Meta Business Suite
  87. Meta Business Help Center
  88. Meta Transparency Center
  89. About Meta
  90. Meta: Instagram platform docs
  91. YouTube Creators
  92. YouTube Official Blog
  93. YouTube Shorts help
  94. How YouTube Works
  95. YouTube Studio
  96. LinkedIn Marketing Solutions
  97. LinkedIn Help
  98. Pinterest Business
  99. Pinterest Business Help
  100. Snapchat for Business
  101. X for Business
  102. Reddit communities
  103. Reddit for Business Help
  104. ASCAP
  105. BMI
  106. SESAC
  107. Global Music Rights
  108. PRS for Music (UK)
  109. PPL (UK)
  110. SOCAN (Canada)
  111. APRA AMCOS (Australia)
  112. GEMA (Germany)
  113. SACEM (France)
  114. SIAE (Italy)
  115. JASRAC (Japan)
  116. IFPI
  117. RIAA
  118. National Music Publishers Association
  119. Harry Fox Agency
  120. SoundExchange
  121. Music Reports
  122. Epidemic Sound
  123. Artlist
  124. Soundstripe
  125. PremiumBeat
  126. AudioJungle
  127. Free Music Archive
  128. Creative Commons
  129. Incompetech
  130. FTC: advertising and marketing
  131. FTC: disclosures 101
  132. FTC: endorsement guides
  133. FTC: consumer reviews rule
  134. FTC: advertising FAQs
  135. US Copyright Office
  136. US Copyright Office: DMCA
  137. US Copyright Office: music FAQ
  138. US Copyright Office: fair use FAQ
  139. USPTO: trademarks
  140. UK Advertising Standards Authority
  141. ACCC (Australia)
  142. Competition Bureau Canada
  143. GDPR overview
  144. California Consumer Privacy Act
  145. COPPA
  146. FTC: children’s privacy
  147. W3C Web Accessibility Initiative
  148. W3C: WCAG
  149. W3C: captions
  150. W3C: making audio and video accessible
  151. ADA.gov
  152. WebAIM
  153. Epilepsy Foundation
  154. Pew Research: internet and technology
  155. DataReportal
  156. US Census Bureau
  157. US Bureau of Labor Statistics
  158. Interactive Advertising Bureau
  159. Think with Google
  160. Google Trends
  161. Nielsen insights
  162. Schema.org: VideoObject
  163. Schema.org: SocialMediaPosting
  164. Schema.org: MusicRecording
  165. Schema.org: HowTo
  166. Schema.org: FAQPage
  167. Schema.org: Organization
  168. Google: video best practices
  169. Google: video structured data
  170. CapCut
  171. Adobe Premiere Rush
  172. DaVinci Resolve
  173. Canva
  174. Descript
  175. VEED
  176. Kapwing
  177. Otter.ai
  178. Later
  179. Buffer
  180. Hootsuite
  181. Sprout Social
  182. Google Analytics
  183. Google Search Console
  184. Google Analytics developer docs
  185. GA4: events and conversions
  186. Matomo
  187. Plausible Analytics
  188. Similarweb
  189. UK Information Commissioner’s Office
  190. Office of the Privacy Commissioner of Canada
  191. Australian OAIC
  192. European Data Protection Board
  193. EU data protection
  194. EU Digital Services Act
  195. Ofcom
  196. FCC
  197. AIGA
  198. Nielsen Norman Group
  199. Smashing Magazine
  200. web.dev
  201. MDN: web media
  202. MDN: the video element
  203. ISO 21001 (reference)
  204. Buma/Stemra (Netherlands)
  205. STIM (Sweden)
  206. Teosto (Finland)
  207. Koda (Denmark)
  208. TONO (Norway)
  209. IMRO (Ireland)
  210. SGAE (Spain)
  211. ZAiKS (Poland)
  212. KOMCA (South Korea)
  213. MCSC (China)
  214. CISAC
  215. World Intellectual Property Organization
  216. TikTok: creating videos
  217. TikTok: exploring videos
  218. TikTok: privacy settings
  219. TikTok: growing your audience
  220. TikTok Creator Academy
  221. TikTok Effect House
  222. TikTok for small business
  223. Instagram: Reels help
  224. YouTube: Shorts best practice
  225. How YouTube recommends
  226. Pinterest Predicts
  227. Snapchat for Business
  228. Hootsuite blog
  229. Social Media Examiner
  230. Marketing Week
  231. Adweek
  232. Google Search Essentials — SEO starter guide
  233. Google: creating helpful, reliable, people-first content
  234. Google: intro to structured data
  235. Google: LocalBusiness structured data
  236. Google: FAQPage structured data
  237. Google: Article structured data
  238. Google: Product structured data
  239. Google: title links in search results
  240. Google: control your snippets
  241. Google: robots.txt introduction
  242. Google: sitemaps overview
  243. Google: consolidate duplicate URLs
  244. Google: redirects and Search
  245. Google: JavaScript SEO basics
  246. Google: multi-regional and multilingual sites
  247. Google Search Central Blog
  248. Google: get started with Search Console
  249. Google: how local search results are determined
  250. Google Business Profile: prohibited and restricted content
  251. Google Business Profile: address and service area guidelines
  252. Google Business Profile: review policy
  253. Google Business Profile: add or edit categories
  254. web.dev: Core Web Vitals explained
  255. web.dev: Largest Contentful Paint
  256. web.dev: Cumulative Layout Shift
  257. web.dev: Interaction to Next Paint
  258. Google PageSpeed Insights
  259. Google Rich Results Test
  260. Google Search Console
  261. W3C Markup Validation Service
  262. Schema.org: LocalBusiness type
  263. Schema.org: Service type
  264. Schema.org: FAQPage type
  265. Schema.org: HowTo type
  266. W3C: WCAG 2.2 quick reference
  267. US Census Bureau QuickFacts: New Jersey
  268. US Census Bureau: American Community Survey
  269. US Census: Statistics of US Businesses
  270. Bureau of Labor Statistics: New Jersey data
  271. BLS: Occupational Employment and Wage Statistics
  272. NJ Department of Labor: labor market information
  273. New Jersey Business Action Center
  274. US Small Business Administration: New Jersey district
  275. USA.gov: business resources

Want this done for your site?We build and maintain the search, content and paid programmes described on this page.

Get a free proposal

Get a free marketing proposal

Tell us what you are trying to grow and we will come back with a plan, not a pitch deck. Same-day reply on weekdays.

Privacy Preferences
When you visit our website, it may store information through your browser from specific services, usually in form of cookies. Here you can change your privacy preferences. Please note that blocking some types of cookies may impact your experience on our website and the services we offer.
Contact Us