Skip to content
← All writings

· 5 min read · Joshua Oas

What are llms.txt and llms-full.txt, and why is everyone adding them?

  • AI
  • SEO
  • Astro

More and more people find information by asking an AI assistant instead of typing into a search box. Ask ChatGPT, Claude or Gemini about a library, a product or a company, and there’s a good chance it fetches a few web pages in the background, reads them and summarizes the answer.

The problem: websites are built for people, not for language models. That’s where llms.txt comes in, and why you’re suddenly seeing it on so many sites.

The problem it solves

When an AI assistant reads your page, it doesn’t see the design. It sees the raw page, and a typical modern page is mostly not content:

  • navigation, footers and cookie banners,
  • scripts, tracking code and inline styles,
  • content that only appears after JavaScript runs (which some fetchers never run),
  • the same header and footer repeated on every page.

Language models also have a limited context window, which is the amount of text they can consider at once. If most of what they fetch is markup and boilerplate, the useful part gets crowded out or cut off. The result is answers that are vague, out of date, or confidently wrong about your product.

Search engines solved a version of this decades ago with robots.txt and sitemap.xml. Those files tell crawlers what they may visit and what exists. Neither one tells an AI what’s important or where the clean version of the content is.

What llms.txt is

llms.txt was proposed by Jeremy Howard of Answer.AI in September 2024, and the spec got a v2 update in 2026 after two years of real-world use. It’s a Markdown file at the root of your site (/llms.txt) with a simple structure:

# Site or project name

> A one or two sentence summary.

Optional paragraphs with any important context.

## Docs

- [Getting started](https://example.com/start.md): Install and first steps
- [API reference](https://example.com/api.md): Every endpoint

## Optional

- [Changelog](https://example.com/changelog.md)

Only the H1 title is required. The blockquote is a summary, the H2 sections are lists of links with short notes, and a section called Optional holds links an AI can skip if it’s short on space.

The key idea is that the file stays small. It’s a map, not the territory. The detail lives behind the links, and an assistant fetches only what it needs for the question it’s answering.

Clean Markdown versions of pages

The spec also recommends offering a plain Markdown version of each page at the same URL with .md added. On this site, /about has /about.md, and every project and post has one too. Each HTML page points to its Markdown version with a link tag:

<link rel="alternate" type="text/markdown" href="/about.md" />

Markdown is exactly the format language models handle best: headings, lists and links, with none of the layout noise.

And llms-full.txt?

llms-full.txt isn’t part of the official spec, but it’s become a common companion, popularized by documentation platforms like Mintlify. Instead of links, it contains the full text of the site in a single Markdown file.

They serve different jobs:

llms.txt llms-full.txt
What’s in it A curated index of links Every page’s full text
Size Small enough to fit in context Can be large
Best for An assistant answering one question Loading a whole site or docs set into a coding assistant or project

Developers often paste an llms-full.txt into a coding assistant so it understands an entire library at once. For a small site like this one, the whole thing is only about 75 KB.

Who’s using it?

According to the spec’s site, thousands of sites now publish one, including the developer docs for OpenAI, Anthropic and Google. Platforms like Mintlify, GitBook, Wix and WordPress SEO plugins such as Yoast can generate it automatically. The spec also notes that Chrome’s Lighthouse checks for one as part of its agentic browsing audits.

The intended use is at inference time: an assistant reads it on demand while helping someone, rather than it being a training-data feed.

What it isn’t

There’s a lot of hype around this, so a few honest caveats:

  • It’s not a ranking signal. No major search engine has said it uses llms.txt to rank pages. Normal SEO (good content, fast pages, clear structure) still matters most.
  • It doesn’t control crawling. If you want to allow or block AI crawlers (like GPTBot or ClaudeBot), that’s still robots.txt.
  • No assistant is guaranteed to read it. Support varies by tool and changes often. Think of it as making your site easy to read if an agent looks, not as a switch you flip.
  • It’s public. Anything in it can be read by anyone. Don’t put private information in it.

How I added it to this site

This site is built with Astro, and all its content lives in content collections, so I didn’t want to maintain these files by hand. They’re generated at build time:

// src/pages/llms.txt.ts
import { buildLlmsTxt, markdownResponse } from '@/lib/llms';
export const GET = async () => markdownResponse(await buildLlmsTxt());

The builder reads every published page, project and post, and writes:

  • /llms.txt: the site summary, then sections for Pages, Projects and Writings, each a link to the Markdown version with its excerpt. Legal pages, the RSS feed and the full-content file go under Optional.
  • /llms-full.txt: every page, project and post as Markdown, each preceded by a comment with its source URL.
  • A .md route for every page, including the page-builder pages, whose sections get converted into headings, lists and links.

Because it’s generated:

  • drafts never show up (including this post, until I publish it),
  • new posts and projects appear automatically on the next build,
  • nothing gets out of date.

I also made sure it only includes what’s already public on the site. No email or phone number, and nothing from the CMS that isn’t shown on a page.

Should you add one?

If you run documentation, a product site or anything people ask AI assistants about, yes. It takes an hour or less, it costs nothing, and it makes your content easier for agents to read accurately. For a personal site, it’s a nice-to-have.

If you’re on a static site generator, generate it from your content instead of writing it by hand. Then forget about it.

You can see mine at /llms.txt and /llms-full.txt.