SEMPITE Research

Original Research on AI Search Visibility

We run live studies against the AI systems people actually use (Google’s AI Overview, ChatGPT and the assistants behind them) and publish what we find. Real queries, recorded answers, named sources. All studies are released under CC BY 4.0 and free to cite.

How AI decides who to recommend

Who gets named in an AI answer, who gets cited as the source, and how little of it goes to the brands being discussed.

Measurement Study · September 2026

Can AI See Your Images? We Checked 2,000 Websites. The Most Visual Ones Scored Worst.

We crawled the images on 2,181 ecommerce, service, and personal/portfolio sites and read each the way a non-rendering AI crawler does. The sites built on images describe them to a machine the worst, and almost nobody labels AI-generated images (or can detect them).

44%of portfolio sites: most images have no usable alt text ~1%carry any image provenance / AI-disclosure metadata 1,082reachable sites' images measured
Measurement Study · September 2026

Half the Stores an AI Agent Can Read Aren’t Ready to Sell to One.

We checked 1,600 Shopify storefronts for AI-shopping-agent readiness. Of the ones an agent can even read, a third serve a blank JavaScript-only product page, 19 in 20 hide their reviews, and only half are transaction-ready.

35%show an agent a blank product page 94.5%expose no agent-readable reviews 52%are transaction-ready
Claims Audit · September 2026

We Fact-Checked 25 Viral AI-Shopping Stats. 11 Failed.

We pulled 130 repeated statistics from 27 sources on AI shopping and ecommerce AEO, then stress-tested the 25 most-cited with three independent adversarial checks each. The “3.1× more citations from schema” stat every AEO deck opens with failed 0–3.

130claims pulled from 27 sources 44%of the most-cited failed (11 of 25) 3.1×schema myth, refuted 0–3
Census · August 2026

We Scored 307 SEO Agencies With Our Own AI-Readiness Rubric

The agencies winning Google’s “SEO agency” results across 30 US metros, each run through our public 24-check rubric. Median 82/100, and then the failure rates start.

1 in 5has no structured data at all 51.5%have no llms.txt 96%allow GPTBot, nobody's blocking AI
Census · August 2026

Only 1 in 9 SEO Agencies Publishes Its Prices

307 ranking agencies crawled for pricing pages and verifiable plan prices, with budget-form decoys excluded by manual review. The industry’s answer to “how much?” is a phone call.

88.6%publish no verifiable prices $1,000median published plan /mo $450median cheapest plan /mo
Cohort Study · August 2026

Time to Visibility: What Five Sites We Operate Did, Day by Day

How long does SEO actually take? Five sites with daily Search Console tracking, measured from first impression to first page-one ranking, first click, and first 100-impression week. Real dates, no guesses.

13median days to first non-brand page-one 30median days to first click 2/5sites clickless after 6+ weeks
Category Study · August 2026

Who Google’s AI Recommends in Sports Nutrition

40 buying-intent queries across energy drinks, pre-workout, creatine and protein. 30 brands scored against every AI Overview answer, every cited source logged.

95%of answers cite the same 10 domains 5.1%of citations go to brand-owned sites 8/30brands never mentioned
Playbook · August 2026

How to Get Cited by the 3 Publishers That Control Supplement AI Answers

Forbes Health, Healthline and Garage Gym Reviews carry 77% of Google’s AI supplement answers, and all three publish how they choose. We read the criteria. They want the same eight things, and none of them is buyable.

77%of AI answers via these 3 publishers 3/3publish their scoring criteria $0you can pay to be ranked
Quarterly edition · September 2026

AI Visibility Index

Whether ChatGPT and Google’s own AI Mode name the businesses Google ranks in its local top three, across 50 U.S. markets. Open data and code.

95%of Google top-3 businesses never named by ChatGPT 68%never named by Google’s own AI Mode

llms.txt and machine-readable access

The file the web adopted to tell AI systems what to read. We fetched every one we could find and opened them.

Ecosystem Audit · August 2026

We Audited 1,563 llms.txt Files. One in Six Is Already Broken.

Every llms.txt file listed in the ecosystem’s public directories, fetched and classified, then 19,039 of the links inside them validated. The web’s instruction manual for AI is quietly rotting.

15.7%of files broken at file level 18.2%of sites link to a dead page 54sites bot-block their own bot file
Ecosystem Audit, Part 2 · August 2026

1 in 8 llms.txt Files Is Empty. The Biggest Has 11,137 Links.

We opened all 1,318 working llms.txt files and counted what’s inside. The format meant to be a curated map is, in practice, either empty, a tidy list, or a full-site firehose, with nothing enforcing the difference.

13%of working files are empty 30links in the median file 11,137links in the largest
Ecosystem Snapshot · August 2026

98% of Shopify Stores Serve the Same llms.txt. It’s Not a Content Map.

We fetched the llms.txt of 10,858 Shopify stores. The file most of them serve is byte-identical, auto-injected, and doesn’t map content. It’s a checkout manifest most merchants don’t know they have.

6,016stores, one identical file 89.7%auto-generated boilerplate 1.5%a genuine content map
Cross-Study · August 2026

The Brands With llms.txt Files Are the Ones AI Ignores

We crossed our two studies: who publishes an llms.txt against who Google’s AI actually cites in sports nutrition. It’s backwards: the sources AI reads skip the file, and the brands it ignores all auto-serve the same commerce template.

1/10top publishers publish any llms.txt 0/17brands with a real content map 7brands serving the same template

Local and service businesses

Where AI answers actually appear in local search, what people search for before they choose a business, and who gets cited.

Measurement Study · September 2026

We Crawled 1,416 Real Estate Agent Websites. Only 1 in 5 Is Ready for AI.

We harvested 1,416 independent agent and brokerage sites across ~55 US markets and read each the way a non-rendering AI crawler does. Most render fine for a human and show an AI assistant no listings, no reviews, and barely an identity. Of the 650 with a listings page, 30.5% serve no property data in HTML; not one of 1,097 sites uses listing schema.

1 in 5clears the AI-ready bar (19.9%) 30.5%show an AI crawler no listings 95.4%expose no machine-readable reviews
Verified Briefing · September 2026

Real Estate Is the Least AI-Answered Industry. The Portals Already Own What’s Next.

We fan-out reviewed real estate SEO, GEO and AEO, then adversarially fact-checked the 25 most load-bearing claims. Real estate triggers a Google AI Overview in just 4.48% of searches, the lowest of any industry, but Zillow, Redfin and Realtor.com are already inside ChatGPT while ~70% of AI crawlers can’t even read a JavaScript-rendered listing.

4.48%AI-Overview rate, lowest of any industry ~70%of AI crawlers can’t read a JS listing 3portals in ChatGPT; independent agents, none
Category Study · August 2026

Google’s AI Skips Local Car Repair Searches. It Owns the Questions That Come First.

102 automotive queries across 12 metros. Google wrote an AI answer for 6.9% of local repair searches, and 90% of the advice questions a driver asks before picking a shop. Those answers are assembled from independent shop blogs and Reddit, with no publisher gatekeeper at all.

6.9%of local repair searches get an AI answer 90%of advice questions do 70:20independent shops vs dealerships cited
Demand Study · August 2026

One Search Term Outweighs Every Mechanical Car Repair in America, Twice Over

We mapped 22.1M monthly US searches across 20 kinds of auto shop. Car washes and tires swamp everything, engine work is almost unsearched, and the priciest clicks sit in the categories nobody looks for. 83% of it says “near me.”

22.1Mmonthly searches mapped 83.1%carry “near me” 2.1×car wash vs all mechanical repair

Products and niche brands

How small brands win crowded product categories: demand maps, ranking concentration, and where the next challenger gets in.

Measurement Study · September 2026

963 Shopify Stores Use 30,622 Different Tags. 700 Steam Games Use 394.

Two catalogues, one primitive, opposite authorship. Steam players tag games from a shared vocabulary; every Shopify merchant invents their own. The median store shares just 16.7% of its tags with any other store, and one store in 1,028 carries every attribute an AI shopping agent reads.

30,622tags across 963 stores 16.7%of a store’s tags shared with any other 1/1,028stores fully readable by an AI agent
Measurement Study · September 2026

72% of Shopify Stores Have No Wishlist. A Third of the Rest Can’t Tell You Anything Came Back.

We sampled 30,000 domains, measured 1,028 Shopify storefronts, then re-checked every wishlist we thought we had found. Installing a wishlist and offering one turn out to be different things, and twelve stores paid for one and switched it off.

28.1%have a save control that works 34.9%of those can never say an item returned 19.4%make you log in before you can save
Demand Study · August 2026

The $369 Dementia Phone That Wins Google With One Blog Post

We mapped the 150 highest-volume Google rankings held by the senior-phone niche leader. 80 of them, 56% of the ranked search volume, ride on a single comparison article that reviews its own competitors, beating brands with five times its search gravity.

80/150rankings on one blog post 56%of ranked volume on that page 5:1rival’s brand demand, RAZ still wins the category

People and personal brands

What AI systems say about individuals who are real but not famous, and what they invent when they do not know.

Category Study · August 2026

We Logged Every Source in 95 AI Book Recommendations. Author Websites Got Zero.

16 reader book-discovery queries, every AI Overview citation recorded and classified. Reddit, BookTube, Goodreads, the book press and small blogs carry the answers: authors’ own sites appear nowhere.

11/16reader queries got an AI answer 0/95citations to author-owned sites 18%of citations to small book blogs
Category Study · August 2026

Google’s AI Knows This Working Actress in Detail. She Has No Wikipedia Page.

20 casting queries, 20 AI answers. IMDb is cited in 80% of them and Backstage in 5%. And an actress with no encyclopedia entry is described accurately, sourced from the website she controls. The barrier isn’t fame, it’s legibility.

20/20casting queries answered 80%of answers cite IMDb 5%cite Backstage
Cross-Surface Study · August 2026

Decline, Fabricate, or Retrieve: Three AI Systems, One Actor, Three Different Answers

We asked ChatGPT, Gemini and Google’s AI Overview about the same 13 performers on the same day. ChatGPT knew all 6 famous ones and refused on all 7 others without inventing anything. Gemini invented a pageant title and a telenovela role. Google looked it up and got it right.

6/6famous actors ChatGPT got right 0/7working actors it would describe 12/13answered by retrieval instead

Put this to work

The studies describe the landscape. These explain what to do about it.

Questions about methodology, data requests, or a category you want studied next: get in touch.

ES