Home › Research › llms.txt and machine-readable access
SEMPITE Research · Cross-Study · August 2026 · By Jacob Gerrish
The Brands With llms.txt Files Are the Ones AI Ignores
We already knew two things about sports nutrition: Google’s AI cites editorial publishers for 95% of its supplement answers, and llms.txt files across the web are quietly rotting. So we crossed the two lists, who publishes an llms.txt against who the AI actually reads. The result is exactly backwards. The sites AI listens to don’t bother with the file. The brands it ignores all have one, and it’s not even the kind of file the format was invented for.
The sources AI actually reads don’t publish llms.txt
Across 39 sports-nutrition AI Overviews, ten editorial domains appeared in 95% of answers: three of them (Forbes, Healthline, Garage Gym Reviews) in 77%. We checked whether each of those ten publishes an llms.txt. Nine do not:
The one exception, tigerfitness.com, is also the least-cited of the ten (7.7% of answers). The publications that dominate AI answers reached that position without ever telling an AI how to read them. Their leverage is editorial authority and structured review content, not a manifest file.
The brands all have one, and it’s a checkout manifest, not a content map
We ran the same check on the 17 brands scored in the study. Not one publishes the kind of llms.txt the format was designed for, a curated map of your best explanatory content. What we found instead:
- 7 serve a byte-for-byte identical template (RAW Nutrition, Nutricost, REDCON1, RYSE, NutraBio, Bloom and Xwerks all publish a file titled “Agent Instructions) how AI agents can interact with our online store.” It’s Shopify’s agentic-commerce manifest (the Universal Commerce Protocol / shop.app checkout skill), injected automatically. It tells an AI how to add to cart and check out, not why to recommend the brand in the first place.
- The other 10 don’t have a working file at all: Optimum Nutrition’s /llms.txt is a redirect loop that dead-ends on a Cloudflare error page; Transparent Labs serves a 0-byte file; Ghost, Red Bull and Guayakí return their homepage HTML (the “looks alive, is useless” failure mode); Celsius and Bucked Up 404; Monster and Reign bot-block it with a 403; and BUM Energy has no cleanly-resolving canonical domain to host one.
The mismatch in one sentence: the brands optimized the file for an AI that buys, when the actual problem, proven by the same study, is that AI never mentions them. A checkout manifest is worthless if the assistant recommends Transparent Labs from a Forbes article and never routes a shopper to your store to begin with.
Why this is backwards, and what actually moves the needle
llms.txt got sold to brands as the way to “show up in AI.” This category shows the opposite: brand-owned domains earn only 5.1% of citation slots regardless of whether they publish the file, and the publishers who own the other 95% skipped it entirely. Publishing a manifest does not buy you a citation. Being the kind of source an AI trusts does.
For a supplement brand, that means the work is upstream of the file: independent lab results and certifications a reviewer can cite, comparison and “best of” content that publishers actually reference, structured Product and Review schema, and relationships with the specific gatekeeper that owns your aisle. As the parent study showed, that is a different publication for pre-workout, creatine, protein and energy drinks. An auto-generated checkout file is not a strategy. It’s a checkbox a platform ticked for you.
Frequently asked questions
Do the sites Google's AI cites publish an llms.txt?
Mostly not. Of the ten editorial domains that appear in 95% of sports nutrition AI answers, nine have no llms.txt at all. The one exception, tigerfitness.com, is also the least-cited of the ten at 7.7% of answers.
Do the supplement brands have llms.txt files?
All 17 brands scored in the study have something at that URL, but 0 of 17 publish a genuine content map. Seven serve a byte-identical Shopify commerce manifest, and the other ten serve broken, empty or HTML responses.
What is the Shopify file the brands are serving?
An auto-injected Agent Instructions file for Shopify's Universal Commerce Protocol and shop.app checkout skill. It tells an AI how to add to cart and check out, not why to recommend the brand.
What share of AI citations go to brand-owned domains?
5.1% of citation slots. The rest go to third-party sources, which is why a manifest on the brand's own site changes little.
What actually earns a place in AI answers, then?
Editorial authority and structured review coverage on the sites the AI already reads. The publications that dominate AI answers reached that position without ever publishing an llms.txt.
Want to know what AI assistants actually say about your brand, and which publishers decide it?
Run a free AI visibility check