Home › Research › How AI decides who to recommend
SEMPITE Research · Claims Audit · September 2026
We Fact-Checked 25 Viral AI-Shopping Stats. 11 Failed — Including the One Every AEO Deck Opens With.
The AI-commerce gold rush runs on repeated statistics. We pulled 130 of them from 27 sources, stress-tested the 25 most-cited with three independent adversarial checks each, and kept only the claims that survived a majority trying to knock them down.
The stat every AEO pitch opens with does not hold up
If you have been sold answer-engine optimisation for an online store, you have probably seen the line: pages with structured data get 3.1× more citations in AI Overviews. It is in decks, in agency landing pages, in LinkedIn carousels. It is the load-bearing number under a lot of invoices.
It failed our verification 0–3 — all three independent checks refuted it. The problem is not that schema is worthless. The problem is that the 3.1× figure describes a correlation between pages that happen to rank and pages that happen to carry schema, then dresses it up as a controlled uplift you can buy. At least one independent test found no citation or ranking benefit from adding schema on its own. Correlation got sold as a lever.
This matters because it is exactly the kind of claim an honest AI-visibility firm should be killing, not repeating. Schema is table stakes for eligibility — most AI-cited pages carry it — but bolting it onto a thin product page will not multiply your citations. If someone quotes you 3.1×, ask them for the controlled test. There isn’t one.
What failed
Eleven of the twenty-five most-repeated claims did not survive. The pattern is consistent: the more precise and the more flattering the number, the less likely it was to hold.
Correlational, presented as causal. Independent testing found no standalone uplift from schema.
Precise failure rates with no traceable methodology. The directional “most stores fail” survives; the decimal points do not.
Single-vendor blog figure stated as an industry constant. Net-margin band was not reproducible across sources.
Fixed dollar CACs quoted as universal. Real CAC depends on margin, AOV and repeat rate — see what held up.
Survey self-report of a hypothetical, presented as measured conversion lift.
The “win on schema alone” inference was not supported by the underlying data.
Also killed: a blanket “average ecommerce CAC is $68–$84, up 40% since 2023” (0–3), the flat “3:1 LTV:CAC is the minimum” rule stated as a law (1–2), a fixed “80–90% fail, 10–20% profitable” when quoted as a precise census (0–3, though the loose version survives), and two more source-specific figures. Full tally ships with the dataset.
What held up — and how confident we are
Fourteen claims survived. These are the ones worth building decisions on, with the honest confidence attached.
The demand shift is real and fast. This is the headline that should replace the myths.
$50 CAC on a $25 thin-margin supplement is underwater; $200 CAC on a $90, 75%-margin, repeat-purchase beauty product is healthy.
The one channel with genuine momentum, corroborated across sources.
Survives as an association: schema is common among cited pages. It does not license the 3.1× causal claim above. Table stakes, not a multiplier.
Agentic commerce is a discovery layer today, not a buying agent. Optimise to be recommended, not to be auto-bought.
Also surviving: Mastercard’s forecast of large AI-driven consumer spend by 2030 (3–0), the loose “most dropshipping stores fail in year one” heuristic (2–1), and Google search CAC sitting in a $50–$130 band with rising CPCs (2–1). Confidence ratings are published per claim.
What this actually means if you sell online
Strip out the dead stats and a clean picture remains. AI shopping is real and growing fast — half of US shoppers already use it. But it is a recommendation and discovery layer, not a checkout robot, and the thing that gets you recommended is not a single schema tag you can buy in an afternoon.
Structured data is necessary but not sufficient: you need it to be eligible, and then the actual levers are the boring, holistic ones — clear product information an answer engine can quote, real review corpus, consistent entity signals across the web, and content that answers the questions shoppers ask AI in the words they ask them. Anyone selling you a guaranteed-citations schema package is selling the 3.1× myth.
That is the difference between chasing a number and earning answer-engine share. It is also, bluntly, why we published the claims that make our own category look overhyped: the firms worth hiring are the ones killing bad stats, not reprinting them.
Frequently asked questions
Does adding schema markup get your pages more citations in AI answers?
The widely-quoted figure that structured data earns 3.1× more AI-Overview citations did not survive verification (refuted 0–3) because it describes correlation, not a controlled uplift, and at least one independent test found no benefit from adding schema alone. What did hold up is that most AI-cited pages happen to carry schema: 65% of Google AI Mode and 71% of ChatGPT-cited pages. Schema is table stakes for eligibility, not a lever that multiplies citations on its own.
How many people actually shop with AI in 2026?
51% of US consumers reported using AI for online shopping in 2025, up from 38%. That survived. But willingness to let AI complete the purchase collapses at checkout, from roughly a quarter or a third happy to research down to about 8% comfortable delegating the buy. AI in 2026 is a discovery and recommendation layer, not a checkout agent.
Is the dropshipping failure rate really 90%?
A loose 80–90% first-year failure rate survived as a directional heuristic, not a precise measurement. The stricter and scarier variants — only 1–5% ever succeed, only 1.5% clear $50k/month — were all refuted 0–3. Most stores fail; the exact percentage is not knowable from public data; anyone quoting a success rate to two decimals is guessing.
What CAC can a dropshipping or ecommerce store actually afford?
There is no universal number. The claim that survived cleanly is that affordable CAC is a function of gross margin, average order value and repeat rate. A $50 CAC on a $25 supplement is underwater; a $200 CAC on a $90 beauty product at 75% margin with repeats is healthy. The fixed CAC dollar ranges various sources quote for Meta and TikTok were mostly refuted.
How did SEMPITE verify these claims?
We extracted 130 falsifiable claims from 27 published sources, took the 25 most-cited, and put each through three independent adversarial checks whose job was to refute it against its source and the wider record. A claim was discarded only if at least two of three refuted it. 14 survived, 11 were killed. This is an LLM-adjudicated audit rather than human peer review, so we publish every claim, its sources and its vote tally for anyone to re-check.
Method, in short
We ran a fan-out search across six angles (dropshipping economics, platform trajectories, ad economics, AI answer engines, agentic commerce, and a contrarian demand-validation pass), fetched 27 sources and extracted 130 falsifiable claims. The 25 most-cited were each cross-examined by three independent adversarial verifiers instructed to refute the claim against its cited source and the wider public record. A claim was discarded only on a majority refute (2 of 3 or 3 of 3).
Known limits, stated plainly: sources skew to industry and vendor blogs, so this audits what the market repeats rather than a controlled ground truth; verification is LLM-adjudicated, not human peer review, which is why we publish every claim, source and vote so you can check our work; and CAC and GMV figures go stale within a quarter or two. Surviving a refutation attempt means a claim is defensible today, not eternally true.
We help online stores earn answer-engine share the honest way — no guaranteed-citation myths.
AI Visibility for eCommerce