SEMPITE Research

Original Research on AI Search Visibility

We run live studies against the AI systems people actually use (Google’s AI Overview, ChatGPT and the assistants behind them) and publish what we find. Real queries, recorded answers, named sources. All studies are released under CC BY 4.0 and free to cite.

How AI decides who to recommend

Who gets named in an AI answer, who gets cited as the source, and how little of it goes to the brands being discussed.

Census · August 2026

We Scored 307 SEO Agencies With Our Own AI-Readiness Rubric

The agencies winning Google’s “SEO agency” results across 30 US metros, each run through our public 24-check rubric. Median 82/100, and then the failure rates start.

1 in 5has no structured data at all 51.5%have no llms.txt 96%allow GPTBot, nobody's blocking AI
Census · August 2026

Only 1 in 9 SEO Agencies Publishes Its Prices

307 ranking agencies crawled for pricing pages and verifiable plan prices, with budget-form decoys excluded by manual review. The industry’s answer to “how much?” is a phone call.

88.6%publish no verifiable prices $1,000median published plan /mo $450median cheapest plan /mo
Cohort Study · August 2026

Time to Visibility: What Five Sites We Operate Did, Day by Day

How long does SEO actually take? Five sites with daily Search Console tracking, measured from first impression to first page-one ranking, first click, and first 100-impression week. Real dates, no guesses.

13median days to first non-brand page-one 30median days to first click 2/5sites clickless after 6+ weeks
Category Study · August 2026

Who Google’s AI Recommends in Sports Nutrition

40 buying-intent queries across energy drinks, pre-workout, creatine and protein. 30 brands scored against every AI Overview answer, every cited source logged.

95%of answers cite the same 10 domains 5.1%of citations go to brand-owned sites 8/30brands never mentioned
Playbook · August 2026

How to Get Cited by the 3 Publishers That Control Supplement AI Answers

Forbes Health, Healthline and Garage Gym Reviews carry 77% of Google’s AI supplement answers, and all three publish how they choose. We read the criteria. They want the same eight things, and none of them is buyable.

77%of AI answers via these 3 publishers 3/3publish their scoring criteria $0you can pay to be ranked
Monthly Index · Updated August 2026

AI Visibility Index

Month-over-month tracking of how often businesses ranking in Google’s top three local results are recommended, or ignored, by ChatGPT.

43%of Google top-3 businesses never mentioned by ChatGPT +11ptsvs July, the gap is widening

llms.txt and machine-readable access

The file the web adopted to tell AI systems what to read. We fetched every one we could find and opened them.

Ecosystem Audit · August 2026

We Audited 1,563 llms.txt Files. One in Six Is Already Broken.

Every llms.txt file listed in the ecosystem’s public directories, fetched and classified, then 19,039 of the links inside them validated. The web’s instruction manual for AI is quietly rotting.

15.7%of files broken at file level 18.2%of sites link to a dead page 54sites bot-block their own bot file
Ecosystem Audit, Part 2 · August 2026

1 in 8 llms.txt Files Is Empty. The Biggest Has 11,137 Links.

We opened all 1,318 working llms.txt files and counted what’s inside. The format meant to be a curated map is, in practice, either empty, a tidy list, or a full-site firehose, with nothing enforcing the difference.

13%of working files are empty 30links in the median file 11,137links in the largest
Ecosystem Snapshot · August 2026

98% of Shopify Stores Serve the Same llms.txt. It’s Not a Content Map.

We fetched the llms.txt of 10,858 Shopify stores. The file most of them serve is byte-identical, auto-injected, and doesn’t map content. It’s a checkout manifest most merchants don’t know they have.

6,016stores, one identical file 89.7%auto-generated boilerplate 1.5%a genuine content map
Cross-Study · August 2026

The Brands With llms.txt Files Are the Ones AI Ignores

We crossed our two studies: who publishes an llms.txt against who Google’s AI actually cites in sports nutrition. It’s backwards: the sources AI reads skip the file, and the brands it ignores all auto-serve the same commerce template.

1/10top publishers publish any llms.txt 0/17brands with a real content map 7brands serving the same template

Local and service businesses

Where AI answers actually appear in local search, what people search for before they choose a business, and who gets cited.

Category Study · August 2026

Google’s AI Skips Local Car Repair Searches. It Owns the Questions That Come First.

102 automotive queries across 12 metros. Google wrote an AI answer for 6.9% of local repair searches, and 90% of the advice questions a driver asks before picking a shop. Those answers are assembled from independent shop blogs and Reddit, with no publisher gatekeeper at all.

6.9%of local repair searches get an AI answer 90%of advice questions do 70:20independent shops vs dealerships cited
Demand Study · August 2026

One Search Term Outweighs Every Mechanical Car Repair in America, Twice Over

We mapped 22.1M monthly US searches across 20 kinds of auto shop. Car washes and tires swamp everything, engine work is almost unsearched, and the priciest clicks sit in the categories nobody looks for. 83% of it says “near me.”

22.1Mmonthly searches mapped 83.1%carry “near me” 2.1×car wash vs all mechanical repair

Products and niche brands

How small brands win crowded product categories: demand maps, ranking concentration, and where the next challenger gets in.

Demand Study · August 2026

The $369 Dementia Phone That Wins Google With One Blog Post

We mapped the 150 highest-volume Google rankings held by the senior-phone niche leader. 80 of them, 56% of the ranked search volume, ride on a single comparison article that reviews its own competitors, beating brands with five times its search gravity.

80/150rankings on one blog post 56%of ranked volume on that page 5:1rival’s brand demand, RAZ still wins the category

People and personal brands

What AI systems say about individuals who are real but not famous, and what they invent when they do not know.

Category Study · August 2026

We Logged Every Source in 95 AI Book Recommendations. Author Websites Got Zero.

16 reader book-discovery queries, every AI Overview citation recorded and classified. Reddit, BookTube, Goodreads, the book press and small blogs carry the answers: authors’ own sites appear nowhere.

11/16reader queries got an AI answer 0/95citations to author-owned sites 18%of citations to small book blogs
Category Study · August 2026

Google’s AI Knows This Working Actress in Detail. She Has No Wikipedia Page.

20 casting queries, 20 AI answers. IMDb is cited in 80% of them and Backstage in 5%. And an actress with no encyclopedia entry is described accurately, sourced from the website she controls. The barrier isn’t fame, it’s legibility.

20/20casting queries answered 80%of answers cite IMDb 5%cite Backstage
Cross-Surface Study · August 2026

Decline, Fabricate, or Retrieve: Three AI Systems, One Actor, Three Different Answers

We asked ChatGPT, Gemini and Google’s AI Overview about the same 13 performers on the same day. ChatGPT knew all 6 famous ones and refused on all 7 others without inventing anything. Gemini invented a pageant title and a telenovela role. Google looked it up and got it right.

6/6famous actors ChatGPT got right 0/7working actors it would describe 12/13answered by retrieval instead

Put this to work

The studies describe the landscape. These explain what to do about it.

Questions about methodology, data requests, or a category you want studied next: get in touch.