Skip to main content
Retail Operations

Returns, and what the data actually tells you

Returns arrive weeks after the order that caused them, reason codes are the weakest field in your database, and the product that looks like a best seller on gross revenue can be losing money on every unit. Here is how to read the data honestly.

The first number everyone looks at is the wrong one

Almost every returns report divides returns received this month by orders shipped this month. That number is not a return rate. It is a ratio between two things that happened at different times, and it moves for reasons that have nothing to do with returns. Run a promotion in October and November's ratio spikes, because October's orders are coming back against November's smaller order base. Grow thirty percent and the ratio falls, flattering you exactly while the absolute cost is rising. The window between purchase and return commonly runs ten to forty-five days depending on the category and the policy, which means the true return rate for last month cannot be known for another month or more.

The fix costs nothing conceptually and is a real change in how the report is built. Cohort by order date. For orders placed in a given week, what fraction has come back so far, and how does that curve mature? You will discover that your return curve has a shape: a steep section in the first two weeks, a long shallow tail out to the policy boundary, and usually a small bump right at the deadline. Once you know that curve you can estimate a fresh cohort's final rate from its first fourteen days, which is what lets you spot a problem in three weeks instead of three months.

Everything else in this article assumes you have made that change first. Analysis built on the received-over-shipped ratio will produce conclusions that reverse when the growth rate does.

You are probably here because

  • Returns went up last quarter and nobody can say whether it was the products, the mix or the growth
  • Somebody suggested tightening the returns window and nobody knows what that would cost
  • A top-ten product by revenue turned out to be unprofitable once returns were counted
  • The reason codes say “changed my mind” on forty percent of returns and that is all anyone knows

The first is a cohorting problem, covered above. The second and third need honest return costing, which is the next section. The fourth is the free-text section, and it is where most of the actionable information in your returns data is sitting unread.

What a return actually costs

Most businesses carry one blended number for the cost of a return, usually one somebody estimated years ago, and it is almost always too low because it stops at freight. Build it up properly once and the rest of the analysis becomes possible.

Inbound freight. Whatever the prepaid label costs, typically five to fifteen dollars domestically depending on size and service.

Receiving and inspection labour. Somebody opens the parcel, checks the item against the order, grades it, and decides its disposition. Three to eight dollars per unit is a fair range, and it goes higher for anything that must be tested, re-packed or re-labelled.

Disposition loss. This is the big one and the one that varies. Sellable-as-new means you lose only the handling. Open-box means a discount of twenty to forty percent. Damaged or worn means liquidation at cents on the dollar, or disposal with a fee attached.

The original outbound cost you do not get back. Shipping out, payment processing on the original sale, and pick-and-pack labour are spent whether or not the item comes home.

The inventory that was not sellable while it travelled. A unit in transit and awaiting grading is a unit you owned and could not sell, and on a seasonal item that time is the whole ballgame. If your third-party warehouse takes eight to fifteen days to process a return back to sellable, that is a real cost hiding in an operations metric nobody reports to finance.

Add those up on a forty-dollar apparel item and the total lands somewhere between fifteen and thirty dollars, which on typical retail margins means one return can consume the profit from two or three sales of the same product. That arithmetic is the whole reason this data is worth analysing.

One return on a forty-dollar item can eat the margin from two or three sales of it. That arithmetic is why a return rate belongs next to the revenue number, not in an operations report.

Reason codes are the weakest field in your database

The dropdown is answered by a person who wants a refund and wants the form to end. They pick the first plausible option, or the safest one. “Changed my mind” and “no longer needed” absorb everything, including genuine quality defects, because a customer selecting “defective” often expects to be asked for evidence and would rather not be.

Three sources are better, in ascending order of usefulness.

The free-text box. Nearly always present, nearly never read. Thirty comments on a single SKU will tell you more than a year of aggregated code counts, and reading them takes an hour. This is the cheapest analysis in this entire article and almost nobody does it, so if you only act on one paragraph here, make it this one.

The warehouse grade. What the receiving team recorded when they opened the box. If sixty percent of one product's returns arrive graded as damaged, the customer's reason code is not the story; the packaging or the carrier is.

The pattern relative to a baseline. A twenty-eight percent return rate on a dress means nothing on its own. A twenty-eight percent rate against a category baseline of thirty-four percent is a good product. The same rate against a category baseline of twelve percent is a defect with a marketing budget behind it. Every return number in every report should be expressed against the baseline for its category, and against the same SKU's own history.

Language models make the free-text step much cheaper than it used to be, and this is one of the genuinely good uses for them: classify a few thousand comments into a stable taxonomy you defined, keep the original text next to the label, and spot-check a sample by hand. It is a two-day job, not a project. Do not let it replace reading the raw comments for your worst twenty products, though, because the specific sentences are where the fix lives.

The metric that changes decisions: contribution net of returns

Gross revenue per SKU is on every dashboard. Contribution net of returns is on almost none, and it is the only figure that should drive a buying or merchandising decision.

For a SKU over a period: units sold, times contribution per unit, minus units returned times the full return cost, minus the disposition loss on those units. Divide by units sold to get contribution per unit shipped. Then rank the catalog by it.

The ranking will not match your revenue ranking, and the gaps are the point. There is usually a product in the top twenty by revenue that sits in the bottom quartile on net contribution, and typically it is a colour or a size rather than a whole style. Everyone in the business has a feeling about that item. Nobody has the number, so the discussion never resolves. Producing the number resolves it in an afternoon.

What the pattern looks likeMost likely causeWhere to look nextThe fix that usually works
One size returns far more than the othersGrading or spec error on that sizeMeasurement chart against actual garmentsCorrect the size chart, add real measurements to the page
Rate jumps on units received after a dateSupplier lot or process changeReceipt records, lot codes, warehouse gradesQuarantine the lot, raise it with the supplier
One colour returns twice the style averagePhotography does not match the productThe free text, and the photo under real lightReshoot; add a swatch and a lifestyle image
Rate rose after a page redesignSpecifications lost in the new layoutPage diff, and time-on-page before the changePut the dimensions and materials back above the fold
Returns arrive graded damagedPackaging or carrier handlingReceiving photos, carrier and laneChange the box or the lane before blaming the product
High rate concentrated in one channelDifferent expectations set by that listingThe listing copy and images on that channelAlign specifications across every channel

Fit is a different problem from quality, and it is fixable

In apparel and footwear, returns run far above the rest of retail, commonly in the twenty-five to forty percent range against a whole-market online figure usually cited somewhere in the mid-teens to twenty percent. The great majority of that gap is fit, not quality, and fit is unusually tractable because the cause is almost always a mismatch between what the page said and what arrived.

The interventions that work are dull: real garment measurements rather than a generic size chart, a measured comparison to a common reference item, photographs on more than one body, a note where a style runs small, and reviews surfaced with the reviewer's own size next to the comment. None of that is a model. All of it is content work, and on a bad SKU it can move the return rate by several points within a season.

Then there is bracketing: the customer orders three sizes intending to keep one. It shows up unmistakably in the data as multiple sizes of one style in a single order. Before you try to stop it, decide whether you want to. For some businesses, bracketing is how a customer buys at all when they cannot try things on, and the customers who bracket are frequently among the highest-spending you have. The right response is usually not to prevent it but to price and plan for it, and to reduce the need for it with better fit information. Attacking it with policy is a good way to lose a cohort you would rather keep.

Where returns work pays off — our ranking by effort against return

Cohorting by order date so the numbers mean something
94
Reading the free text on the worst twenty products
88
Net-of-returns contribution per SKU on every report
82
Fit content: real measurements, better photos, sized reviews
76
Lot-level tracing from return back to receipt
58
Predicting at checkout whether an order will be returned
22

Our ranking, not a measurement of your business. The last row is low because a good prediction rarely comes with an action you are willing to take at checkout.

The supplier signal hiding in your returns

Returns are the only place a customer inspects your product for you, at scale, for free. Most retailers never connect that inspection back to the goods that produced it.

The connection needs one field: which receipt, lot or production run the returned unit came from. If your warehouse records a lot code at receiving and the return record keeps it, you can compute a return rate per lot, and a defect that entered on one production run becomes visible in weeks instead of at the end of a season. Without that field you are averaging good product and bad product together and concluding the SKU is mediocre.

If you cannot get lot codes, receipt dates are a usable approximation. Bucket units by which receipt they most likely came from, using the order dates and your stock movements. It is inexact and it still finds step changes, which is the pattern you are hunting.

Serial returners, and the trap in acting on them

A small fraction of customers, often between one and three percent, generate a share of returns far out of proportion to their numbers. It is tempting to identify them and restrict them. Before you do, compute their net contribution, because the same list very often contains some of your highest-spending, highest-frequency customers. A customer who orders eighteen times a year and returns a third of it may still be worth several times an average customer, and a policy that pushes her away is expensive in a way that will never show up in the returns report that motivated it.

Split the population properly. High return rate with high net contribution is a good customer with a shopping style. High return rate with negative net contribution across a long history is a different case, and even there the first response should be service rather than restriction: better fit guidance, a note about the item, a proactive size suggestion.

Genuine abuse is a narrow category, and it is a detection problem rather than an analytics one. Empty boxes, swapped items, worn goods returned as new. The answer is inspection at receiving with photographs, and a small manual review queue. Do not build a model to catch it. Build a record that lets a person confirm what came back, and keep a rate on how often the box did not contain what the customer said.

Returnless refunds are arithmetic, not generosity

For low-value items, getting the goods back costs more than the goods are worth. Compare the cost to recover, which is inbound freight plus inspection plus the handling, against expected recovery, which is the resale price times the probability the item comes back sellable. Below roughly fifteen to twenty-five dollars of item value, recovery frequently loses, and it loses harder for anything hygiene-related, perishable, or heavily personalised, where the item cannot be resold at all.

Compute the threshold per category rather than picking one number, publish it internally, and apply it as a rule in the returns flow. Two guardrails: cap how often a single customer can receive a returnless refund, and exclude items above a value where the abuse risk changes the calculation. Otherwise this is one of the few places where the cheapest option and the best customer experience are the same option.

Analysis Note

Keep the returned unit joined to the original order line

The single most common data defect we find in returns work is a return table that records SKU, date and amount but not which order line it reversed. Without that join you cannot compute a cohort curve, you cannot attribute a return to a channel, a promotion, a discount level or a lot, and you cannot compute net contribution correctly. It is usually one column and a small backfill, and it unlocks most of the analysis in this article. Fix it before anything else.

Send the returns table and we will tell you what is in it.

Twelve months of order lines and return records, plus your product export, to contact@precisionfederal.com. You get back the cohort curve, the ten SKUs costing you the most net of returns, and what we would fix first. One business day, written, no charge.

contact@precisionfederal.com

The mistakes we find most often

  • Return rate computed as returns received over orders shipped in the same month
  • Returns not joined to the original order line, so nothing can be attributed to anything
  • One blended return cost applied to a nine-dollar accessory and a four-hundred-dollar coat
  • Reason codes reported as fact while the free-text field has never been opened
  • Return rates compared across categories with no baseline, so apparel always looks broken
  • Refunds recorded at the parent product, hiding the one size or colour causing it
  • Policy tightened as the first move, before anyone costed what it would save or lose

Three weeks that get you most of it

Returns analysis, in order

1
Join returns to order lines; build the cohort curve by order week and by category
Days 1–3
2
Cost a return properly, per category, including disposition and days out of stock
Days 4–5
3
Rank the catalog by contribution net of returns; find the gaps against the revenue ranking
Days 6–8
4
Read the free text on the twenty worst; classify the rest and keep the raw text attached
Days 9–12
5
Fix pages, size charts, photos and packaging on those twenty; note the date you changed each
Days 13–15
6
Stand up the weekly report on cohorts and net contribution; re-measure the twenty in eight weeks
Week 4 onward

Step five is the one that matters and the one that gets postponed, because it is merchandising work rather than analysis. Write down the date of every page change, or you will not be able to tell in December whether the improvement came from the fix or from the season.

When you do not need software for this

Under a few hundred orders a month, none of this needs a system. Export the last year to a spreadsheet, add a column for order week, compute the cohort rate, sort by net contribution, and read every comment. It is a day of work, it will find the same things a build would find, and the honest recommendation is to do that first regardless of size, because it tells you whether there is enough here to justify anything more.

Build when the analysis needs to be repeated weekly by someone who will not write SQL, when returns span several channels and warehouses that record things differently, or when the returns process itself needs automating rather than the reporting. Those are real reasons. “We should have a returns dashboard” is not, and a dashboard built on the wrong ratio is worse than nothing because it produces confident wrong answers on a schedule.

Before you change the returns policy

  • Every return joined to the order line it reverses
  • Rates cohorted by order date, with a maturation curve per category
  • Return cost built up per category, including disposition and days out of stock
  • Net-of-returns contribution computed at variant level, not parent
  • Rates always shown against a category baseline and the SKU's own history
  • Free text read for the twenty worst products by return dollars
  • Lot or receipt traced from the return back to the goods
  • Serial returners assessed on net contribution before any restriction
  • A returnless-refund threshold computed per category and written down
  • The cost and the revenue effect of any policy change estimated before it ships

Bottom line

Returns data is better evidence than most retailers realise and it is almost always read wrong. Cohort by order date so the rate means something. Cost a return honestly, including the days the unit was not sellable. Rank the catalog by what it contributes net of returns, and act on the products where that ranking disagrees with the revenue ranking. Read the comments. Trace the lot. And treat a policy change as the last resort rather than the first, because it lowers returns by lowering sales, and the report that motivated it will not show you that half.

Frequently asked questions

What is a normal return rate?

It depends almost entirely on category. Online retail overall is commonly cited in the mid-teens to around twenty percent; apparel and footwear frequently run twenty-five to forty; consumables, beauty and hard goods are often under ten. Comparing your rate to a headline market figure tells you nothing useful. Compare each product to its own category and to its own history.

Should we shorten the returns window to cut returns?

It will cut returns and it will also cut orders, and the second effect is invisible in the report that suggested the change. Before touching the policy, cost the returns properly and fix the products driving them. If you do run the experiment, measure revenue per visitor and repeat purchase rate, not the return rate, because the return rate will improve whether or not the change was good for you.

Is it worth predicting which orders will be returned?

Rarely, because the prediction usually comes with no action you are willing to take. You are not going to block the order or charge that customer more. The exception is a soft intervention at the moment of choosing: a size suggestion, a fit note, a prompt when someone adds three sizes of one style. That is worth building, and it needs far less modelling than a return prediction.

How long before we can tell whether a product page fix worked?

Roughly your return curve's maturation window plus enough orders to see a difference, so commonly six to ten weeks for a moderately popular item. Record the change date, compare cohorts before and after, and be careful about seasonality: a fix made in October and measured in December is confounded with the gifting season.

Do we need a returns platform, or can our current systems do this?

Most platforms handle the customer-facing workflow well and the analysis poorly, because they hold the return and not the original order economics. The usual answer is to keep the platform for the workflow and do the analysis where your order, cost and inventory data already live. What you need there is one clean join between the return and the order line.

1 business day response

Want to know which products are losing money once returns are counted?

Send a year of order lines and return records with your product export. We will build the cohort curve, rank the catalog on contribution net of returns, and tell you what we would fix first, or do the work with your team. Email bo@precisionfederal.com.

Email an engineerCapabilitiesMore insights →
Retail AnalyticsReverse LogisticsUnit EconomicsData Engineering