Skip to content
Utilities2026-10-05Begin4 min read

llms.txt, Explained: A Site Manual for LLMs — and Whether Yours Deserves One

The problem llms.txt solves

An LLM processing a website faces a dilemma: it crawls the homepage, wades through navigation, ads, and footer noise, and can't tell what's worth reading or how to understand it. Traditional SEO infrastructure doesn't help — robots.txt only governs "may you crawl", sitemap only lists "what pages exist", and no file answers "what is this site's knowledge structure".

llms.txt (proposed September 2024 by Jeremy Howard of Answer.AI) fills the gap: a markdown file at the site root that serves as a manual for large models. The path is fixed at /llms.txt, and the format is human-readable markdown — a deliberate choice, since LLMs digest markdown far better than XML.

The spec: three-part structure

# Site Name

> One sentence on what this site is and who it serves (blockquote summary)

## Essential
- [Core page](https://example.com/path): one sentence on why this page is worth reading
- [Another core page](https://example.com/path2): one sentence

## Optional
- [Secondary content](https://example.com/path3): one sentence

Three design points:

  1. Every link carries a one-sentence annotation — the essential difference from a bare link list; the annotation helps the model decide whether to expand that link
  2. Tiers (essential / optional) — model context is finite; priority information saves its compute directly
  3. Markdown, not XML — compared to the machine format of sitemap.xml, llms.txt is "prose written for a model to read"

Division of labor with existing SEO files

FileQuestion answeredConsumer
robots.txtMay you crawlAll crawlers
sitemap.xmlWhat pages existSearch engines
llms.txtWhat's worth reading, and how to frame itLLMs and AI applications

All three coexist, each covering its own lane. Shipping llms.txt doesn't exempt you from robots configuration, or vice versa (allow/block decisions in the AI crawler guide).

The honest controversies

To be clear, llms.txt is no silver bullet:

  • Google's public stance is reserved — the search team has stated it is not used for ranking; AI vendors have made no unified commitment on whether or how they consume it
  • Adoption is growing but the absolute numbers remain small (some major sites in, most of the long tail out)
  • It does not single-handedly produce rankings or citations — without good content, the shiniest manual is worthless

So why ship one? Three pragmatic reasons: the cost is near zero (half an hour); there are no side effects (if nobody consumes it, you lose nothing); and when a citation decision happens, it improves your hit quality — a clear manual is a tiebreaker when a model chooses among candidate sources. The full GEO strategy is in the introduction.

Practice: generate it, don't maintain it

Hand-maintained llms.txt inevitably rots (the moment a page changes, the manual is stale). This site's approach is build-script generation: the file is derived from the tool registry and blog index, refreshed automatically on every deploy — llms.txt always matches site structure at zero marginal cost.

The advice for independent site owners is the same: if your site structure lives in a single source of truth (even a YAML config), generate llms.txt from it. If not, at least put "sync llms.txt" into the publishing checklist — an outdated manual is worse than none.

Implementer's note

Two details from writing the generator: first, the annotation wording matters more than you'd think — "online tool" carries zero information for a model, while "a JWT decoder where input never leaves your browser" is what makes the model recall you in the right context. Second, llms.txt should also expose entrances to markdown versions of content pages where the site supports them — models digest markdown sources at far lower cost than HTML, the cheapest concession a content site can make to AI.