Skip to main content
Merchandising & Personalization

Recommendations for a store with a small catalog

With eight hundred products, most pairs of items have never been bought together even once. That single fact decides which surfaces are worth building, which rules beat a model, and how long you will have to wait to know whether any of it worked.

Start with the arithmetic of your own data

Before anyone picks an algorithm, do this multiplication. Say the store carries eight hundred sellable items. That gives about three hundred and twenty thousand possible pairs. Now count multi-item orders: if you did forty thousand orders last year and, as is common for a specialty store, sixty to seventy percent of them contained exactly one item, you have somewhere around fourteen thousand baskets that can tell you anything about pairing at all. Those baskets produce maybe twenty to thirty thousand pair observations, and they concentrate heavily on the same few dozen popular items. The plain result is that well over ninety percent of pairs in your catalog have never been observed together, and most of the ones that have were seen once.

That is not a reason to give up. It is the design constraint. It tells you that a pure co-purchase approach will be confident about your forty best sellers and silent about everything else, which is exactly backwards from where the merchandising opportunity usually sits. It tells you that product attributes, which you already have and which cover the whole catalog evenly, will carry more weight than behavior for most items. And it tells you that a merchandiser writing a hundred and fifty hand-picked pairs in an afternoon is a genuinely competitive system, not a stopgap.

You are probably here because

  • An app added a “you may also like” strip and nobody can tell whether it made any money
  • The recommendations under a product are four near-identical versions of that product
  • Someone bought a mattress and the site spent three weeks recommending mattresses
  • A vendor quoted personalization and the traffic numbers do not look big enough to test it

The section on complements and substitutes covers the second and third. The measurement section covers the first and fourth, and it is the one worth reading twice, because on a small store the measurement is harder than the model.

Recommendations are a placement decision before they are a model decision

The same algorithm put in two different slots produces completely different economics, and stores usually build the slot with the worst economics first because it is the most visible. Six surfaces are worth thinking about, and they want different things.

The product page strip. The customer is evaluating one item and has not decided. What helps here is usually a substitute: the same thing in another colour, a cheaper version, the next size up. It reduces bounce and it moves customers toward a purchase, but be clear-eyed that some of what it earns is revenue you would have got anyway from a person who was going to buy something.

The cart. The decision is already made and the customer is in a spending frame. This is where complements belong: the blade for the razor, the case for the phone, the filter for the pitcher. Attach rates of two to eight percent on a cart module are ordinary, and the incremental revenue is much easier to believe than a product page strip because the base order was already committed.

Post-purchase email. Underrated and usually the highest-return surface on a small store. A note three days after delivery suggesting the accessory, or a replenishment reminder timed to actual consumption, converts at rates onsite widgets do not reach, costs nothing to send, and can be evaluated with a clean holdout because you control who receives it.

Search with no results. Somebody typed a word and got an empty page. Anything relevant beats nothing, and this is the cheapest win in the whole list. Pull the query log for zero-result searches; on most small stores it is a short and very revealing document.

The homepage. The weakest surface for a small catalog, because a returning customer already knows what you sell and a new one has no history. Fresh arrivals and best sellers, updated by a person, are hard to beat here.

The order confirmation and the account page. Reorder prompts for consumables. Not glamorous, and for a store selling coffee, filters, supplements or pet food it can be worth more than everything else combined.

Somebody bought a mattress and the site spent three weeks recommending mattresses. Purchase should suppress a category, not amplify it.

Complements and substitutes are not the same list, and mixing them costs money

This is the single most common defect we see, and it comes from behavior data being blind to the difference. A co-view model learns that people who looked at the charcoal sweater also looked at the oatmeal one, which is true and useless in the cart, where the customer already has a sweater. A co-purchase model on a sparse catalog often cannot say anything at all about the pairing you care about.

Product attributes solve this cleanly and cheaply. Same category and adjacent price band means substitute. Different category with a declared compatibility, or a documented consumable relationship, means complement. Most stores already hold enough attribute data to build both lists, and the ones that do not can fix it with an afternoon of tagging on the top two hundred items.

Then use them by slot: substitutes on the product page, complements in the cart and in the follow-up email. Getting that one rule right typically does more for revenue than any model swap, and it takes a day.

Two adjacent traps. Variants are not products. If a shirt exists in nine sizes and five colours, recommending four of its own variants is a broken widget that will look fine in every offline metric. Roll variants to the parent before scoring and again before rendering. Durable goods should be suppressed after purchase. Somebody who bought a mattress is not a mattress lead; they are a pillow and a mattress protector lead. Maintain a per-category rule for how long a purchase suppresses the category and how long it boosts the accessories.

SurfaceWhat belongs thereHow to build it firstHow to judge it
Product page stripSubstitutes: colour, size, price stepAttribute similarity within categoryBounce and add-to-cart, and be honest about cannibalization
Cart moduleComplements and consumablesHand-curated pairs on the top 100 sellersAttach rate and incremental order value on a holdout
Post-purchase emailAccessories, then replenishment timed to useCategory rules plus a consumption interval per SKUHoldout revenue per recipient; the cleanest test you have
Zero-result searchAnything relevant at allCategory fallback plus a synonym list from the query logExit rate on the results page
HomepageNew arrivals, best sellers, seasonalA merchandiser and a calendarUsually not worth an experiment at this size

The rules are most of the system, and they are not a compromise

Every serious recommendation deployment ends up with a rules layer on top of whatever scores the items, and on a small catalog that layer is doing most of the work. Write it deliberately instead of discovering it through complaints.

Never recommend an out-of-stock item, or one that ships in six weeks. Never recommend across a category boundary that makes no sense, and keep that list as data rather than in code, because merchandisers need to edit it. Suppress the item already on the page and its own variants. Suppress anything already in the cart. Apply a margin or strategic weighting where you want to, and write down that you did, because a recommender quietly optimized toward high-margin items is a decision the business should make on purpose. Cap how often the same item can appear across a session so the widget does not turn into a single-product billboard.

The one to argue about is the margin weighting. It works, it is legitimate, and it degrades trust if it is heavy-handed. Our starting position is to use it as a tiebreaker among items of similar relevance rather than as a term with real weight, and to look at the recommendation set with fresh eyes once a month.

Effort against likely return — small catalog, our ranking

Curated complements in the cart, top 100 sellers
90
Replenishment email timed to consumption
86
Fixing zero-result search
80
Attribute similarity on the product page
72
Co-purchase model over the whole catalog
48
Per-visitor personalization on the homepage
25

Our judgment for a catalog in the hundreds with a few thousand sessions a day. It inverts above roughly ten thousand items and heavy repeat traffic. The ordering is the point.

Personalization means less than the word suggests at this size

Personalization implies a model of a person, and a model of a person requires repeated observations of that person. For a store where a typical customer buys twice a year and browses anonymously in between, there is almost no person to model. Most sessions are first sessions or effectively first sessions, and the honest version of personalization at this scale is contextual: what is this visitor looking at right now, what did they put in the cart, what did they search for, are they on a phone at eleven at night.

There is one exception, and it is worth building. Customers who have ordered three or more times are a different population with real history, and for them a genuine preference model is both feasible and valuable, particularly in email. Build for that segment specifically rather than pretending the whole audience is knowable.

Measuring it when traffic is thin

This is where small stores get taken advantage of, usually not deliberately. A vendor reports revenue attributed to recommendations, that number is large, and the number is close to meaningless, because it usually counts any order in a session where a recommendation was clicked, including orders for items the customer was going to buy regardless.

The arithmetic is unforgiving. Three thousand sessions a day at a two percent conversion rate is sixty orders a day. Detecting a five percent relative lift in revenue per session against ordinary order-value variance takes weeks and often more than a month, and that is one test. If the roadmap has twelve ideas on it, you cannot test them one at a time in a year.

Three practical answers. Test the big swings, not the small ones: a new surface, a rules change with real teeth, a switch between substitutes and complements in the cart. Prefer revenue per session as the metric over click-through, because click-through on a recommendation widget improves easily by showing the customer things they were already going to buy. And run the email tests first, since a holdout on a send list is clean, cheap and fast, and what you learn there about which pairings work transfers directly onsite.

Where a proper experiment is not affordable, a switchback is workable: run the change on alternating weeks for six to eight weeks and compare. It is weaker than a randomized split and it is far better than an attributed-revenue dashboard.

Measurement Note

Ask any vendor exactly one question

“How is attributed revenue computed, and what happens to an order where the customer clicked a recommendation for an item they did not buy?” If the answer credits the whole order, the number is a session counter with a currency symbol in front of it. Ask instead for a holdout: a randomly selected slice of traffic that sees no recommendations at all, reported weekly. Any serious vendor will agree to it, and the ones that resist are telling you something.

Send us your order lines and we will tell you if this is worth building.

A year of orders with SKUs, a product export with categories and attributes, and your session count, to contact@precisionfederal.com. We will compute the basket sparsity, tell you which surfaces have enough signal behind them, and name what we would do first. One business day, written, no charge.

contact@precisionfederal.com

What we usually find

  • Variants recommending their own siblings, so the strip under a shirt shows four sizes of that shirt
  • Out-of-stock items recommended because the widget reads a catalog export refreshed nightly
  • The same module and the same list on every surface, cart and product page alike
  • Attributed revenue reported with no holdout, and nobody able to say what it means
  • A durable good recommended to the person who just bought one, for weeks
  • Zero-result searches never looked at, though the log names the products customers wanted and could not find
  • A model trained on clicks from a widget that only ever showed best sellers, which then confirms that best sellers work

An order of work that fits a real store

Six weeks, in the order that pays

1
Basket sparsity, single-item order share, zero-result query log, repeat-buyer count
Week 1
2
Attribute clean-up on the top 200 items; declare complement and substitute relationships
Week 2
3
Rules layer and live inventory check; cart module with curated complements
Week 3
4
Post-purchase and replenishment emails with a permanent holdout group
Week 4
5
Attribute similarity on the product page, variants rolled to parent
Week 5
6
Measurement: sitewide holdout, weekly revenue per session, a review the merchandiser runs
Week 6

Notice that no model appears until week five, and that a co-purchase model does not appear at all. On a catalog of this size that is usually the right sequence. If the store grows past a few thousand items and repeat traffic gets dense, revisit it; the arithmetic at the top of this article is the thing that changes.

When you should not build this

If most orders are single-item and always will be because you sell one expensive thing, recommendations are a distraction and the money is in the product page, the photography and the shipping promise. If your catalog is under about a hundred items, a customer can see all of it, and navigation beats recommendation. If search is bad, fix search first: a customer who types a query has told you exactly what they want, and failing that request is more expensive than any missed cross-sell.

And if the honest answer is that a merchandiser who knows the products can write the pairs by hand in two afternoons, do that. It will be better than a model trained on twenty thousand sparse observations, it will be right on the items where the model would be silent, and it costs almost nothing. Build the machinery when the hand-written list stops being maintainable, which happens somewhere around a few thousand items or when the catalog turns over faster than a person can keep up with.

Before you turn it on

  • Variants roll up to the parent product before scoring and before rendering
  • Availability checked live at render, not from a nightly export
  • Substitutes on the product page, complements in the cart and in email
  • Post-purchase suppression rules per category, with a duration written down
  • A permanent holdout, onsite and in email, reported weekly
  • Revenue per session is the headline metric; click-through is diagnostic only
  • Merchandisers can edit the rules and the pairs without an engineer
  • Zero-result searches reviewed monthly by a person who can add products or synonyms
  • Any margin weighting is explicit, bounded, and known to the business

Bottom line

On a small catalog, recommendations are a merchandising system with a little bit of modelling in it, not a modelling system with some merchandising around the edges. The data you have is thin in exactly the places behavior-based methods need it to be thick, and dense in the place attributes cover well. So put substitutes where somebody is choosing, complements where somebody has chosen, curate the top hundred by hand, fix the searches that return nothing, and put a holdout behind all of it so you can tell whether any of it worked. That is a few weeks of work and it beats most of what gets sold as personalization at this size.

Frequently asked questions

How small is too small for product recommendations?

Under about a hundred items, a customer can browse the whole catalog and good navigation does more than any recommender. Between roughly one hundred and a few thousand items, curated pairs plus attribute similarity is the right shape. Behavior-driven models start to earn their keep when the catalog is large enough that people cannot see it all and dense enough that most item pairs have been observed together more than once.

Which placement should we build first?

The cart, with hand-curated complements on your top hundred sellers, and the post-purchase email. Both sit after the purchase decision, so what they earn is easier to believe as incremental, and both can be built in days. The product page strip is the most visible surface and the hardest to prove, which is why it should not be first.

A vendor says recommendations drive twelve percent of our revenue. Is that real?

Usually not as stated. Most attribution credits the entire order to any session where a recommendation was clicked, including items the customer had already decided to buy. The only number worth trusting is a holdout: a random slice of traffic that sees no recommendations, compared on revenue per session. Ask for it before you renew.

Can we do this without enough traffic to run experiments?

Partly. Email tests work at much lower volumes because the holdout is clean and you control the send. Onsite, a switchback design running the change on alternating weeks for six to eight weeks gives a usable read when a randomized split would take too long. Either way, test a few large changes rather than many small ones.

Should recommendations favour higher-margin products?

Mildly, and on purpose. Used as a tiebreaker among items of similar relevance it is legitimate and profitable. Given real weight it starts surfacing items customers do not want, which costs trust and eventually costs conversion. Whatever you choose, make it an explicit parameter that someone reviews, not a quiet term buried in a scoring function.

1 business day response

Not sure whether recommendations would pay on your store?

Send a year of order lines, a product export and your session volume. We will compute how much signal your baskets actually contain, say which surfaces are worth building, and give you the order we would do them in, or build it with you. Email bo@precisionfederal.com.

Email an engineerCapabilitiesMore insights →
MerchandisingRetail AnalyticsExperimentationMachine Learning