Home › Research › Products and niche brands
SEMPITE Research · Measurement Study · September 2026
963 Shopify Stores Use 30,622 Different Tags. 700 Steam Games Use 394.
Two catalogues, one primitive, opposite authorship. On Steam the crowd tags games from a shared vocabulary. On Shopify every merchant invents their own. We measured what that difference actually costs.
Every store builds its own private language
963 Shopify stores in this sample used 30,622 distinct tags between them. The median store shares just 16.7% of its tag vocabulary with any other store in the sample. Five stores out of six use words nobody else uses.
Compare 700 Steam titles, which drew on 394 distinct tags in total. The median title shares 100% of its tags with other titles, because there is nothing else to share: Steam tags come from one controlled vocabulary that every game draws from.
This is why Steam can silo a new release by genre within days of launch and let a shopper browse sideways into games they have never heard of, and why a Shopify tag only filters inside the store that wrote it. A tag with a shared meaning is infrastructure. A tag with a private meaning is a label.
What the median catalogue actually looks like
The typical store puts 4 tags on a product and carries 96 distinct tags across its catalogue. The structure underneath those tags is where it gets expensive.
A namespaced tag is written material:cotton rather than cotton, so a filter app can read the prefix as a facet. Only one store in twenty does this for most of its tags.
Casing drift is the quiet one. 31.1% of stores carry the same tag in more than one capitalisation, such as Application and application. A human reading the admin sees one tag. Every filter app sees two, and the products split between them. The worst store in the sample had 322 of these collisions.
Where crowd tagging concentrates
Steam titles carry up to 20 displayed community tags each, and the vote counts show how much agreement sits behind them. The median title has 3,385 tag votes, and the top tag accounts for only 8.5% of them: agreement is broad rather than dominated by one label.
The 20 most common tags account for 33.8% of all tag applications. The vocabulary has a heavy head and a genuine tail, which is what a working classification system looks like. A Shopify tag list, by contrast, is almost all tail.
The part that matters in 2026
Tags are for humans browsing your store. AI shopping agents do not browse; they read structured feed attributes. We checked every store’s product markup against the 12 core attributes that shopping feeds and agent integrations expect.
Most stores are most of the way there. The median store carries 8 of 12. But exactly one store out of 1,028 carried all twelve. A deep, carefully maintained tag taxonomy sitting on top of an empty structured-data layer is work the machines that increasingly route demand simply cannot see.
Frequently asked questions
How many tags should a product have?
The median Shopify store in this sample puts 4 tags on a product, with the middle half of stores between 1 and 9. But the count matters far less than whether the tags mean anything outside your own store. 4.3% of stores use no tags at all, and the deepest store in the sample averaged 171 per product.
What is a tag namespace and does it matter?
A namespaced tag is written as a prefix and a value, such as 'material:cotton' or 'application:outdoor', so a filter app can read the prefix as a facet name. Only 23.7% of stores use any namespacing and just 5.2% apply it to most of their tags. Without it, tags are a flat pile of strings and adding a new filter dimension means touching every product.
What is tag casing drift?
It is when the same tag exists in more than one capitalisation, such as 'Application' and 'application'. A person reading the admin sees one tag; a filter app sees two, and products split between them. 31.1% of stores in this sample had at least one such collision, and the worst had 322.
Why do Steam and Shopify tags behave so differently?
Authorship. Steam tags are applied by players from a shared controlled vocabulary of a few hundred official tags, so 'Souls-like' means the same thing on every game and the platform can use it for discovery across the whole catalogue. Shopify tags are invented by each merchant for their own store, so 963 stores produced 30,622 different tags and almost none of them are comparable.
Do product tags help AI shopping agents find my products?
Not directly. Agents read structured feed attributes, not merchant tags. 79.1% of stores publish Product JSON-LD, but the median store carries 8 of the 12 core attributes and only one store in 1,028 carried all 12. GTIN appears on 18.1% and MPN on 10.2%. A deep tag taxonomy with an empty structured-data layer is work that agents cannot see.
Method, in short
Seeded random sample of 30,000 domains from the Tranco top 1M, probed at /products.json to identify Shopify storefronts. 1,075 found, 1,028 measured, up to 250 products enumerated per store. The Steam arm sampled 700 titles from the store search frame and read community tags and vote counts from each store page. All proportions carry Wilson score intervals.
Known limits: the Tranco frame is global and traffic-weighted, so this describes trafficked Shopify stores rather than all of them; roughly one Shopify store in nine disables /products.json; Steam displays only a title’s top 20 tags, so tags-per-game is censored and only vocabulary and concentration are measurable; structured-data completeness is read from one product page per store; and the two platforms are not matched samples, so the comparison is structural rather than controlled. Full limitations ship with the dataset.
We run this catalogue audit against individual stores as part of the Brand Audit.
See the Brand Audit