HomeResearchProducts and niche brands

SEMPITE Research · Measurement Study · September 2026

963 Shopify Stores Use 30,622 Different Tags. 700 Steam Games Use 394.

Two catalogues, one primitive, opposite authorship. On Steam the crowd tags games from a shared vocabulary. On Shopify every merchant invents their own. We measured what that difference actually costs.

16.7%of a store’s tags are shared with any other store
5.2%of stores use a tag naming convention
31.1%have tags that differ only by capitalisation
1store in 1,028 is fully readable by an AI shopping agent

Every store builds its own private language

963 Shopify stores in this sample used 30,622 distinct tags between them. The median store shares just 16.7% of its tag vocabulary with any other store in the sample. Five stores out of six use words nobody else uses.

Compare 700 Steam titles, which drew on 394 distinct tags in total. The median title shares 100% of its tags with other titles, because there is nothing else to share: Steam tags come from one controlled vocabulary that every game draws from.

Shopify · 963 stores30,622
Steam · 700 games394

This is why Steam can silo a new release by genre within days of launch and let a shopper browse sideways into games they have never heard of, and why a Shopify tag only filters inside the store that wrote it. A tag with a shared meaning is infrastructure. A tag with a private meaning is a label.

What the median catalogue actually looks like

The typical store puts 4 tags on a product and carries 96 distinct tags across its catalogue. The structure underneath those tags is where it gets expensive.

Use any namespacing23.7%
Mostly namespaced5.2%
Have casing drift31.1%
Store logic in tags10.7%
No tags at all4.3%

A namespaced tag is written material:cotton rather than cotton, so a filter app can read the prefix as a facet. Only one store in twenty does this for most of its tags.

Casing drift is the quiet one. 31.1% of stores carry the same tag in more than one capitalisation, such as Application and application. A human reading the admin sees one tag. Every filter app sees two, and the products split between them. The worst store in the sample had 322 of these collisions.

Where crowd tagging concentrates

Steam titles carry up to 20 displayed community tags each, and the vote counts show how much agreement sits behind them. The median title has 3,385 tag votes, and the top tag accounts for only 8.5% of them: agreement is broad rather than dominated by one label.

Singleplayer 531 games Atmospheric 241 games Action 354 games RPG 226 games Adventure 326 games Story Rich 199 games Multiplayer 271 games Strategy 198 games Indie 251 games Casual 192 games Simulation 244 games Exploration 192 games

The 20 most common tags account for 33.8% of all tag applications. The vocabulary has a heavy head and a genuine tail, which is what a working classification system looks like. A Shopify tag list, by contrast, is almost all tail.

The part that matters in 2026

Tags are for humans browsing your store. AI shopping agents do not browse; they read structured feed attributes. We checked every store’s product markup against the 12 core attributes that shopping feeds and agent integrations expect.

Publish Product JSON-LD79.1%
Carry a brand62.1%
Carry a GTIN18.1%
Carry an MPN10.2%
Carry all 12 attributes0.1%

Most stores are most of the way there. The median store carries 8 of 12. But exactly one store out of 1,028 carried all twelve. A deep, carefully maintained tag taxonomy sitting on top of an empty structured-data layer is work the machines that increasingly route demand simply cannot see.

Frequently asked questions

How many tags should a product have?

The median Shopify store in this sample puts 4 tags on a product, with the middle half of stores between 1 and 9. But the count matters far less than whether the tags mean anything outside your own store. 4.3% of stores use no tags at all, and the deepest store in the sample averaged 171 per product.

What is a tag namespace and does it matter?

A namespaced tag is written as a prefix and a value, such as 'material:cotton' or 'application:outdoor', so a filter app can read the prefix as a facet name. Only 23.7% of stores use any namespacing and just 5.2% apply it to most of their tags. Without it, tags are a flat pile of strings and adding a new filter dimension means touching every product.

What is tag casing drift?

It is when the same tag exists in more than one capitalisation, such as 'Application' and 'application'. A person reading the admin sees one tag; a filter app sees two, and products split between them. 31.1% of stores in this sample had at least one such collision, and the worst had 322.

Why do Steam and Shopify tags behave so differently?

Authorship. Steam tags are applied by players from a shared controlled vocabulary of a few hundred official tags, so 'Souls-like' means the same thing on every game and the platform can use it for discovery across the whole catalogue. Shopify tags are invented by each merchant for their own store, so 963 stores produced 30,622 different tags and almost none of them are comparable.

Do product tags help AI shopping agents find my products?

Not directly. Agents read structured feed attributes, not merchant tags. 79.1% of stores publish Product JSON-LD, but the median store carries 8 of the 12 core attributes and only one store in 1,028 carried all 12. GTIN appears on 18.1% and MPN on 10.2%. A deep tag taxonomy with an empty structured-data layer is work that agents cannot see.

Method, in short

Seeded random sample of 30,000 domains from the Tranco top 1M, probed at /products.json to identify Shopify storefronts. 1,075 found, 1,028 measured, up to 250 products enumerated per store. The Steam arm sampled 700 titles from the store search frame and read community tags and vote counts from each store page. All proportions carry Wilson score intervals.

Known limits: the Tranco frame is global and traffic-weighted, so this describes trafficked Shopify stores rather than all of them; roughly one Shopify store in nine disables /products.json; Steam displays only a title’s top 20 tags, so tags-per-game is censored and only vocabulary and concentration are measurable; structured-data completeness is read from one product page per store; and the two platforms are not matched samples, so the comparison is structural rather than controlled. Full limitations ship with the dataset.

We run this catalogue audit against individual stores as part of the Brand Audit.

See the Brand Audit

SEMPITE Research, September 2026. n = 1,028 Shopify storefronts and 700 Steam titles, collected 13 to 14 September 2026. Raw data and all collection code published under CC BY 4.0.

Related studies

All SEMPITE research →

ES