If you’ve read our guide to what llms.txt is, you know that the standard includes two files: llms.txt (the summary) and llms-full.txt (the comprehensive version). Most people’s instinct is that the full version must be more useful – more content, more detail, more information for AI systems to work with.
The data tells a different story. And it’s one of the more interesting things we’ve discovered since building Getmd.
What we’re measuring
Getmd tracks every request that AI crawlers make to our customers’ endpoints. Every page fetch, every file request, every response code – all logged with the specific crawler’s identity, timestamp, and the path they requested.
This means we have real data on how often AI crawlers request llms.txt versus llms-full.txt across hundreds of websites. Not estimates. Not assumptions based on what we think AI systems should do. Actual request logs from real AI crawlers.
The pattern
Across our customer base, llms.txt consistently receives significantly more requests than llms-full.txt. The summary file gets hammered. The full file gets occasional visits.
This isn’t a marginal difference. In most cases, llms.txt receives five to ten times the request volume of llms-full.txt. Some sites see even more dramatic ratios.
The pattern holds across different industries, different website sizes, and different AI crawler families. ClaudeBot, GPTBot, PerplexityBot – they all show the same preference for the summary file.
Why this makes sense
The instinct that “more information is better” doesn’t account for how AI systems actually work during real-time retrieval.
When someone asks ChatGPT or Perplexity a question, the AI system needs to find relevant sources quickly. It’s not doing a comprehensive research project – it’s answering a question in seconds. The retrieval process looks something like this:
- Identify potentially relevant websites
- Fetch a lightweight summary of each site to confirm relevance
- Identify the specific pages that contain the answer
- Fetch those specific pages
- Extract and synthesise the answer
llms.txt is purpose-built for step 2. It’s concise – typically 500 to 1,000 tokens – and gives the AI system exactly what it needs: who is this website, what do they cover, and where are the pages that matter. The AI system can read it in a fraction of a second and make a relevance decision.
llms-full.txt is built for a different use case. It’s comprehensive – potentially tens of thousands of tokens for a large site. Reading and processing it takes significantly more time and compute. For real-time question-answering, that overhead isn’t justified when the summary file already provides enough information to identify the relevant pages.
Think of it this way: when you’re looking for a restaurant, you check the menu summary outside, not the 40-page recipe book in the kitchen.
Where llms-full.txt does get used
This doesn’t mean llms-full.txt is useless. We do see it getting requested – just less frequently and by different patterns.
Periodic deep crawls. Some AI systems do comprehensive crawls of websites for training data or knowledge base updates. These aren’t real-time retrieval events – they’re scheduled indexing runs. For these, the full file is useful because the system isn’t under time pressure and wants to understand the complete content landscape.
Initial discovery. When an AI system encounters a new website for the first time, it sometimes requests the full file to build a comprehensive understanding before subsequent visits rely on the summary.
Research-mode queries. Some AI systems distinguish between quick-answer mode and deep-research mode. When a user asks for comprehensive analysis rather than a quick answer, the system may fetch more comprehensive source material – including llms-full.txt.
The practical implication: you should have both files. But the one that drives the majority of your AI visibility is llms.txt.
What this means for your strategy
This data has clear practical implications:
Invest most of your effort in llms.txt. This is the file AI systems actually use to make retrieval decisions about your website. Make it excellent. Write a clear, accurate description of your business. Link to your highest-value pages – the ones you most want AI systems to find and cite. Keep the descriptions concise but specific.
Keep it current. Because llms.txt is the primary discovery file, stale information here has outsized impact. If you’ve launched a new product, added a new service line, published a major piece of content, or changed your positioning – your llms.txt should reflect that. This is one of the reasons Getmd manages this file for you – changes to your site are reflected automatically.
Be strategic about what you include. Not every page on your website belongs in llms.txt. Think about the questions people ask AI systems about your industry, and link to the pages that answer those questions. Your pricing page, your core service pages, your pillar content, your case studies – these are the pages that drive AI visibility. Your cookie policy and your sitemap are not.
Don’t neglect llms-full.txt entirely. Create it, keep it reasonably current, and let it serve the periodic deep-crawl use case. But don’t agonise over it the way you should over your summary file.
Use analytics to validate. With Getmd’s analytics, you can see exactly which files are being requested, how often, and by which crawlers. This lets you validate that your llms.txt is being consumed and iterate based on real data rather than assumptions.
The llms.txt quality checklist
Based on what we see performing well across our customer base, here’s what makes a strong llms.txt:
A clear, specific description. Not “we’re a leading provider of innovative solutions.” Instead: “We’re a B2B SaaS platform that provides inventory management for mid-size ecommerce businesses across 14 countries.” Specific, factual, and immediately useful for relevance matching.
Links to 10–20 of your most important pages. Not your entire sitemap – your highest-value pages. Service pages, product pages, pillar guides, pricing, case studies. Each with a one-line description of what the page contains.
Accurate topic coverage. List the topics your website covers so AI systems can quickly match your site to relevant queries. These should map to the actual content on your site, not aspirational positioning.
A link to llms-full.txt. Include a reference to the full file for AI systems that want it, but let the summary carry the primary discovery load.
Current information. This seems obvious, but it’s the most common failure point. If your llms.txt references a service you no longer offer or doesn’t mention a product you launched last month, AI systems are working from an outdated picture of your business.
The bigger picture
The llms.txt vs llms-full.txt data is one of the clearer examples of a broader principle: AI visibility isn’t about giving AI systems everything. It’s about giving them the right things in the right format at the right time.
A concise, well-structured summary outperforms a comprehensive data dump because AI systems – like humans – work more effectively with curated, relevant information than with raw volume.
This principle extends beyond llms.txt to your page content (clean markdown beats bloated HTML), your content structure (answer-first formatting beats buried insights), and your discovery infrastructure (specific, current files beat stale comprehensive ones).
See which files AI crawlers are actually requesting on your site. Start your free trial →
This is part of our series on making your website visible to AI. Also in this series:
- Introducing Getmd: Make Your Website Visible to AI
- What is Markdown and Why Every AI System Prefers It
- What is llms.txt? The Discovery File AI Crawlers Actually Read
- Why Getmd Converts Live and Caches Every 5 Minutes
- The Clean Markdown Problem: Why Most Converters Produce Garbage
- Which AI Bots Are Visiting Your Website? (And Which Aren’t)
- How to Read Your AI Crawler Analytics and Fix What’s Broken
- Getting Started with Getmd: Connect Your Website in 5 Minutes
- Getmd for WordPress, HubSpot, Shopify, and Webflow





