Published Jul 19, 2026

llms.txt won the traffic war and you didn't notice

By Kevin Champlin

llms.txt won the traffic war and you didn't notice

The moment I realized traffic had fundamentally changed shape

Three weeks into deploying Vantage AI (our equities ensemble backed by Claude), I noticed something odd in the application logs: the model was pulling financial data from a client's WordPress site—one that ranked well for "quarterly earnings methodology"—but the request came not from a Googlebot or human browser. It came from Perplexity's crawler.

That site's GA4 dashboard showed zero traffic attribution. Zero sessions. Zero engaged time. But Vantage was getting cited predictions that depended on that page's structured data existing and being indexable by an LLM.

A week later, we audited the client's actual traffic shape. Search referrals looked flat year-over-year. But when we traced where our AI model was pulling answers from, we found it was fetching from 47 distinct pages that Google hadn't sent a single visitor to in six months. Perplexity had. Claude.com had. The structured answer engines had.

That's when I understood: SEO optimization for answer engines is not a future state. It's happening right now. And the teams winning are the ones who stopped thinking about page views and started thinking about structured retrievability.

Why your GA dashboard is lying to you

Google Search Console measures clicks. GA4 measures sessions and conversions. Both are useless for answer-engine readiness because an LLM doesn't click through—it extracts. Perplexity's crawler fetches your page, indexes the schema markup or prose structure, synthesizes it with five other sources, and returns an answer. The user never sees your URL. Your traffic counter sees nothing.

This isn't new. Search engines have been extracting answer boxes, featured snippets, and knowledge panels for years. What's new is the scale and openness: Claude.com, Perplexity, and a dozen other AI chatbots now explicitly crawl sites that publish llms.txt files. No opt-out form. No rate-limit negotiation. If you're crawlable and structured, you're fair game.

For a Fortune 500 apparel brand we worked with, this mattered. Their product guides and care instructions were being cited in Claude.com without a single tracked conversion. The queries were real—tens of thousands monthly—but the model was summarizing their knowledge instead of sending traffic their way. The brand's internal team initially saw this as theft. I saw it as a distribution channel they hadn't optimized for.

llms.txt is the new robots.txt, but nobody's treating it that way

An llms.txt file is a plain-text manifest you place at /.well-known/llms.txt that tells LLM crawlers which of your pages are safe to index for answer extraction. It's optional, but it's also a signal: you're aware of the channel, and you're intentional about how your content flows into it.

The teams I know who are treating it seriously—and there aren't many—are doing three things:

  • Listing high-intent pages explicitly. Not every page deserves to be in llms.txt. FAQs, product specs, technical documentation, and methodology pages should be there. Blog listicles and opinion pieces generally shouldn't. One regulated beverage portfolio client we work with lists only their official serving guidance, nutritional methodology, and product safety documentation. That's 140 pages out of 3,200 on the site.
  • Maintaining clear schema markup. If a page is in llms.txt, its structured data needs to be bulletproof. Schema.org FAQPage, Product, Article, and HowTo markups don't just help Google anymore—they're the primary input for LLM answer synthesis. A bad FAQ schema with mismatched questions and answers doesn't break the SERP, but it completely breaks Claude.com's ability to cite you accurately.
  • Treating answer-engine traffic as a separate conversion funnel. You won't measure it in GA4 sessions. But you can measure it as citations (using a tool like Semrush's Brand Monitoring, or just reading Claude.com transcripts manually for a month). One client went from 0 citations to 180 monthly citations across their top 15 target queries in 90 days. The queries drove zero GA sessions but 340 inbound phone calls in the same window—measured through a separate call-tracking UTM parameter added to pages in llms.txt.

The real risk: you're already being indexed, whether you know it or not

This is the part teams get wrong. Perplexity doesn't ask for permission. Claude.com does check llms.txt, but if you don't publish one, it falls back to robots.txt. If your site isn't explicitly blocked there, you're already being crawled for answer extraction. Your content is already being cited in LLM outputs. You just have zero control over how.

I shipped a diagnostic into our AI Showcase that crawls a site for llms.txt compliance and checks for common schema failures—missing author attribution, broken FAQ markup, rel=canonical conflicts. The first time we ran it on a client's site, we found 23 pages being cited by answer engines that had schema errors bad enough to confuse the LLM's extraction logic. Those pages were generating false citations. The model was quoting them but paraphrasing wrong, which hurt both the client's authority and the end user's trust in the answer.

Most teams don't even know this is happening.

The counterargument I hear, and why it's wrong

"But Kevin, won't this just commoditize my content? If Claude cites it directly, why would anyone visit my site?"

Fair question. And the answer is: sometimes they won't. Some queries are terminal—the user gets their answer and moves on. But most aren't. In our testing with the apparel brand, 34% of phone inbound mentions and product purchases traced back to users who had first encountered the brand in a Claude.com answer, then followed the citation to the site for more detail, pricing, or to place an order. Without the citation, those users would have gone to a competitor's site in Claude's next turn.

The real win is being the source of record, not hiding. If Claude's going to answer the question anyway, you want your structured data to be the cleanest, most authoritative input in its context window.

What to do Monday morning

Audit your llms.txt presence. Check if you have one. If not, create one at /.well-known/llms.txt with a clear list of on-topic pages (200–500 high-signal URLs is typical). Include a Disallow rule for pages you don't want indexed for answers—blog commentary, user-generated content, paywalled sections.

Run a schema audit. Use Google's Rich Results Test on your llms.txt pages. Fix any errors. Schema isn't about Rankings anymore; it's about whether an LLM can extract a coherent answer from your page without hallucinating.

Set up a separate conversion tracking for answer-engine queries. Add a UTM parameter or a data attribute to tracked pages. Use Semrush, Ahrefs, or manual search to find where Claude and Perplexity are citing you. Map those queries to downstream actions. You won't see GA sessions, but you'll see phone calls, form submissions, and purchases.

Treat this as a continuity plan, not an add-on. If search shifts more of its answer volume to LLM outputs—and it will—you're betting your revenue on whether you're cited accurately and first. Control that by controlling your structure.

The traffic that doesn't look like traffic is the only traffic that matters anymore.

At Champlin Enterprises, we've integrated llms.txt compliance and schema validation into WordPress and headless CMS modernizations—particularly for clients whose organic strategy depends on answer-engine discoverability. See how we've applied this to recent engagements.

Free Tool

See exactly what AI costs — across every provider.

MyTokenTracker is a free, multi-provider intelligence platform with live pricing across 100+ models. Compare Claude, GPT-4o, Gemini, and more side-by-side — built for developers evaluating models, teams tracking API spend, and founders building AI-native products who want to stay cost-aware before it becomes a line item worth explaining.