No GUI Research · September 8, 2026 · 1,061 words
97% of llms.txt files never received a single AI bot visit. Three independent studies close the debate on the format. What actually produces results: content structured with attributed statistics, expert quotes, and a direct answer at the top of the document.
Key concepts: llms.txt · Generative Engine Optimization (GEO) · SDSR · RAG · organic AI citation
| Study | Sample | Key finding |
|---|---|---|
| Ahrefs | 137,000 domains | 97% of llms.txt files: zero AI bot visits |
| OtterlyAI | 62,100 bot requests (90 days) | Only 0.1% reached /llms.txt |
| Allmo.ai | Top 50 globally cited domains | 0% had active llms.txt |
llms.txt is a Markdown file at a website’s root designed for language models to read as a site summary. The promise was simple: give the AI a clean guide to your content and it will cite you more.
Jeremy Howard (Answer.AI) proposed it in September 2024 for development tools like Cursor, Windsurf, and Claude Code — coding assistants that need to navigate technical documentation without processing heavy HTML. The format was designed so a developer could tell their IDE “here’s my library documentation in a clean format.”
The problem emerged when the digital marketing world adopted it as a generative AI visibility solution — a use case it was never built for.
Platforms like OpenAI, Google, and Anthropic already have their own HTML extraction pipelines that clean code, remove ads, and automatically extract the main text. Accepting an unverified file that the site owner can freely manipulate would introduce direct manipulation risks that no serious laboratory can tolerate.
97% of analyzed llms.txt files received exactly zero AI bot visits.
Ahrefs audited 137,000 websites with the file active. Of the 3% that received any request, 96% came from infrastructure audit bots like BuiltWith — not generative engines. Real AI tool traffic came almost entirely from code assistants like Claude Code, confirming the format’s original purpose.
Additional finding: no AI bot searched for an llms.txt file on a site that hadn’t previously published one. There is no proactive discovery of the format.
Of 62,100 recorded AI model bot requests, only 84 reached /llms.txt — 0.1% of the total.
OtterlyAI monitored AI bot traffic for three months on domains with the file active. The most revealing result: /llms.txt received three times less traffic than any standard HTML page on the same domain. The “AI-special” file gets less bot attention than an ordinary article.
Only 1 of the top 50 globally cited domains had llms.txt. In the Top 20 most-cited media domains: 0%.
Allmo.ai cross-referenced the domains appearing most frequently in ChatGPT, Perplexity, and Gemini responses against llms.txt presence. The conclusion is direct: there is no statistical correlation between having the file and being cited. The world’s most-referenced domains reached that position without it.
Mark Williams-Cook (Search Engine Journal) demonstrated that the four indicators used to sell llms.txt as effective are empty metrics.
Williams-Cook created cats.txt — a file with completely
fictional data about office cats, including invented feline productivity
statistics. The file passed all four “proof points” that agencies and
tools presented as evidence that llms.txt works:
None of these points prove organic citation preference in real user searches. They only confirm that a plain text file is HTTP-accessible and that an LLM can read what it’s directly pointed to — exactly what any ordinary HTML page would do.
The academic research with real empirical evidence points to two concrete strategies, neither of which involves silent files.
RAG precision: +29.6% in standard systems, +29.8% in multi-hop agent architectures (arXiv:2604.19777)
Transformers have primacy bias: content that appears at the beginning of a document receives greater attention and has higher citation probability. SDSR exploits this by positioning key information before everything else.
Implementation rule: every page begins with (1) a 2-3 sentence executive summary, (2) a list of 3-5 key concepts, (3) the main quantitative data point or table. Development follows after. Never the other way around.
Citation visibility: +30% to +40% (Aggarwal et al., Princeton/ACM SIGKDD 2024, GEO-bench, n=10,000 queries)
Each content section must follow this exact order:
| Technique | Citation visibility increase |
|---|---|
| Attributed source citations | +34.4% |
| Quantitative statistics | +32.1% to +37.0% |
| Direct expert quotes | +29.7% |
| Keyword stuffing (old SEO) | Minimal to negative |
| JSON-LD in isolation | +0.17 (irrelevant) |
The most powerful finding for non-market-leader brands: sites in Google positions 4–5 that applied GEO achieved +115% visibility in generative AI responses. Models prioritize extracted fragment quality, not historical domain authority. That’s GEO’s democratizing effect — and the reason mid-size brands can outrank giants in AI answers.
llms.txt solves a real problem — but for developers documenting software, not for brands trying to appear in ChatGPT or Perplexity.
Time invested in creating and maintaining llms.txt files produces zero measurable impact on organic citation. The same investment applied to structuring existing content with SDSR and the GEO atomic template produces increases of 30% to 40%.
The evidence is definitive. The debate is closed.
Want to know where your brand stands in AI models today?
Sources: Ahrefs Crawl Analysis (n=137,000 domains, 2026) · OtterlyAI Bot Traffic Report (n=62,100 requests, 90 days, 2026) · Allmo.ai Citation Analysis (Top 50 global domains, 2026) · Williams-Cook, M. “The cats.txt Experiment”, Search Engine Journal (2026) · Aggarwal et al. “GEO: Generative Engine Optimization”, Princeton/ACM SIGKDD 2024, arXiv:2311.09735 · arXiv:2604.19777 (SDSR Framework)
What's your AI visibility score?
Drop your URL. We'll generate your AML and run your first diagnostic.
Run my diagnostic →