How llms.txt files follow the format: 438 files counted
Most llms.txt files get the top of the format right and drift below it. Of 438 files named llms.txt in public GitHub repositories, 96.6% open with an H1 and 85.2% follow it with a summary blockquote. But only 30.6% follow the whole structure the llms.txt proposal sets out: in 52.7% of files, at least one section under an H2 heading holds prose, not the list of links the proposal describes.
| Rule from the llms.txt proposal | Files | Share |
|---|---|---|
| Opens with an H1, the only required section | 423 | 96.6% |
| Has exactly one H1 | 394 | 90.0% |
| Has a summary blockquote before the first H2 | 373 | 85.2% |
| Uses no headings but the H1 and its H2s | 354 | 80.8% |
| Keeps every H2 section a list of links | 143 | 32.6% |
| Follows every structural rule above but the blockquote | 134 | 30.6% |
What 438 llms.txt files look like
159 of the files sit at the root of their repository; the rest live in folders such as a documentation site's public directory, where a build publishes them. Almost all of them have structure: 419 (95.7%) divide their content under H2 headings. The median file holds 13 links, and 102 files (23.3%) hold none at all.
The H1 and the blockquote, the part of the format every guide shows first, are where files agree: 96.6% open with the H1 the proposal requires, and 85.2% carry the one-paragraph summary it describes next. 89 files (20.3%) keep a section named Optional.
Where files drift from the format
The proposal describes everything after the summary in two parts: sections of any kind except headings, then sections under H2 headers that are each a "file list", a markdown list whose items are a [name](url) link with optional notes after a colon. Files part ways with it in the second part. Only 143 files (32.6%) keep every H2 section a list of links, 19 of them because they have no H2 sections at all. In 231 (52.7%), at least one H2 section holds paragraphs of prose, which makes the file read like a document with chapters more than a map of links. 84 files (19.2%) also use H3 or deeper headings, or a second H1, which the format does not describe.
None of this stops an agent reading the file: it is markdown either way, and the proposal expects agents to view or search the file and follow its links. What drifts is the predictability the format promises, that a program can find the links by reading the H2 lists alone.
Where the links point
The proposal asks that the links point to content a language model can read easily, and suggests offering a markdown version of each page at the same address with .md appended, or with its extension replaced by .md. Of the 56,557 full web addresses in these files, 6,170 (10.9%) point straight at a .md, .mdx or .txt file. That understates how many reach markdown, because v2 also lets a page point to its markdown version with a link relation, which a count of addresses cannot see.
119 files (27.2%) link mostly by relative path rather than by full address. A relative link is read against wherever the file is fetched from, so it works when the file is served from the site it describes.
Check your own llms.txt
To find H2 sections that are not pure link lists, print the first line in each one that is not a list item opening with a link, skipping indented lines, which continue the item above them. This is the output from a scratch llms.txt with a Docs list (one note wrapped onto a second line, a nested item and an H3), a Pricing section written as prose, and an Optional list:
$ awk '/^#/{s=($0 ~ /^## /)?$0:""; next} s && NF && !/^[-*+] \[/ && !/^[ \t]/{print NR": "s": "$0; s=""}' llms.txt
18: ## Pricing: Acme charges per request, billed monthly.Each line it prints is a section to turn into a list of links, or to move above the first H2, where the format allows any content but headings. It reads every line as text, so a ## line inside a fenced code block also counts as a heading. Real llms.txt files from other projects are on the llms.txt page, most starred first, and how to add one to a site is in how to add an llms.txt file.
How we counted
The registry collects agent markdown from public GitHub repositories and stores every version of each file by the SHA-256 of its content. We took every file named exactly llms.txt it holds, leaving out llms-full.txt and other names, read each one's latest stored version, and checked its bytes against its hash; 4 files with no stored version yet were left out. The proposal is written for files served on websites; these are the files as committed to repositories, many of which a build publishes to a site.
A leading byte-order mark was dropped and headings inside fenced code blocks were ignored. A file opens with an H1 when its first non-blank line is a level-one heading. The blockquote counts when a line beginning with > appears before the first H2. An H2 section runs to the next heading of any level. An indented line in it is read as part of the list item above it, and skipped when no item comes before it. The section is a list of links when every other non-blank line is a list item made of a [name](url) link, optionally followed by a colon and notes; it holds prose when a line of text, outside any code block, is not a list item. A code block fails the list test without counting as prose, and a file with no H2 sections counts as keeping every H2 section a list. A file follows the whole structure when it opens with its only H1, uses no other headings but H2s, and keeps every H2 section a list of links; the blockquote is left out of that test because the proposal requires only the H1. A link ends in markdown when its path ends in .md, .mdx or .txt. The registry crawls the public repositories it has found that hold agent files, not all of GitHub, so read the shares as describing those repositories.
Sources
Read next
- How to add an llms.txt file to a website
- Every llms.txt in the registry, most starred first
- How projects write Cursor rules: 557 files counted
Try it
npx modelranch add anthropics/skills/pdf