llms.txt, Explained: A Site Manual for LLMs — and Whether Yours Deserves One
The problem llms.txt solves
An LLM processing a website faces a dilemma: it crawls the homepage, wades through navigation, ads, and footer noise, and can't tell what's worth reading or how to understand it. Traditional SEO infrastructure doesn't help — robots.txt only governs "may you crawl", sitemap only lists "what pages exist", and no file answers "what is this site's knowledge structure".
llms.txt (proposed September 2024 by Jeremy Howard of Answer.AI) fills the gap: a markdown file at the site root that serves as a manual for large models. The path is fixed at /llms.txt, and the format is human-readable markdown — a deliberate choice, since LLMs digest markdown far better than XML.
The spec: three-part structure
# Site Name
> One sentence on what this site is and who it serves (blockquote summary)
## Essential
- [Core page](https://example.com/path): one sentence on why this page is worth reading
- [Another core page](https://example.com/path2): one sentence
## Optional
- [Secondary content](https://example.com/path3): one sentence
Three design points:
- Every link carries a one-sentence annotation — the essential difference from a bare link list; the annotation helps the model decide whether to expand that link
- Tiers (essential / optional) — model context is finite; priority information saves its compute directly
- Markdown, not XML — compared to the machine format of sitemap.xml, llms.txt is "prose written for a model to read"
Division of labor with existing SEO files
| File | Question answered | Consumer |
|---|---|---|
| robots.txt | May you crawl | All crawlers |
| sitemap.xml | What pages exist | Search engines |
| llms.txt | What's worth reading, and how to frame it | LLMs and AI applications |
All three coexist, each covering its own lane. Shipping llms.txt doesn't exempt you from robots configuration, or vice versa (allow/block decisions in the AI crawler guide).
The honest controversies
To be clear, llms.txt is no silver bullet:
- Google's public stance is reserved — the search team has stated it is not used for ranking; AI vendors have made no unified commitment on whether or how they consume it
- Adoption is growing but the absolute numbers remain small (some major sites in, most of the long tail out)
- It does not single-handedly produce rankings or citations — without good content, the shiniest manual is worthless
So why ship one? Three pragmatic reasons: the cost is near zero (half an hour); there are no side effects (if nobody consumes it, you lose nothing); and when a citation decision happens, it improves your hit quality — a clear manual is a tiebreaker when a model chooses among candidate sources. The full GEO strategy is in the introduction.
Practice: generate it, don't maintain it
Hand-maintained llms.txt inevitably rots (the moment a page changes, the manual is stale). This site's approach is build-script generation: the file is derived from the tool registry and blog index, refreshed automatically on every deploy — llms.txt always matches site structure at zero marginal cost.
The advice for independent site owners is the same: if your site structure lives in a single source of truth (even a YAML config), generate llms.txt from it. If not, at least put "sync llms.txt" into the publishing checklist — an outdated manual is worse than none.
Implementer's note
Two details from writing the generator: first, the annotation wording matters more than you'd think — "online tool" carries zero information for a model, while "a JWT decoder where input never leaves your browser" is what makes the model recall you in the right context. Second, llms.txt should also expose entrances to markdown versions of content pages where the site supports them — models digest markdown sources at far lower cost than HTML, the cheapest concession a content site can make to AI.
Related reading
- What is GEO? Generative engine optimization, introduced
- AI crawlers: allow or block, with robots.txt recipes
Related Tools
Related Articles
What Is GEO? Generative Engine Optimization: The New Traffic Frontier of AI Search
GEO (Generative Engine Optimization) is to AI search what SEO is to Google — when ChatGPT, Perplexity, and Google's AI Overviews answer questions directly, users stop clicking links, and being cited by AI becomes the new traffic entry. This post breaks down the four essential differences from SEO, the four core practices (crawlability, structure, citability, fact density), and real GEO data from a working tool site.
Allow or Block AI Crawlers: GPTBot, ClaudeBot, PerplexityBot — a robots.txt Decision Guide
35%+ of top-1,000 sites now handle AI crawlers in robots.txt, but there is no one-size answer — and training crawlers are separate agents from search crawlers (blocking GPTBot does not remove you from ChatGPT Search). A crawler-by-crawler usage table, a three-type decision framework, copy-paste robots.txt rules, and how to measure the payoff after allowing.
The Complete Guide to PDF Tools: Merge, Split, Compress, and Convert — All in Your Browser
Everything you need to work with PDF files without uploading them: when to merge vs split, how PDF compression actually works, converting PDF to JPG and back, and why browser-local processing matters for contracts and ID documents.