Use caseReporting and analyticsWorkflow automation
What your returns are telling you: return reasons turned into product and size-guide fixes
Groups return reasons, reviews and support tickets per SKU, variant and batch into a weekly issue list, so fit and quality problems get fixed at the source.
A blueprint, not a client story. The business described is illustrative; the architecture, integrations and trade-offs are real, and this is how I would build it. By Ergini, .
The short version
A weekly job that reads every return reason, review and product-related support ticket, ties each one to its order line, variant and production batch, and sorts the comments into a fixed list of fit, quality and expectation themes. Code counts them against each product's normal level, and the product team gets a short digest in Slack with counts, batch splits and customers' own words. It reads Shopify, Loop or AfterShip Returns, the helpdesk, review platforms and the ERP. People decide every fix.
- Best for
- Fashion, footwear and outdoor brands where a few categories drive most returns and the reason-code report says 'size' every month.
- Connects to
- Shopify, Loop or AfterShip Returns, Gorgias or Zendesk, Reviews (Okendo, Yotpo, Trustpilot), ERP or PIM, Slack
- The AI does
- Classifies short, multilingual comments into themes with the supporting phrase, names new clusters, and writes digest lines from numbers computed by code.
- People do
- Accept new themes, check the flagged batches, and decide on size-guide changes, photo retakes and supplier claims.
- Built as
- AI Workflow Automation, usually $3.5K - $12K
Every returns report says 'size'
Take an outdoor clothing brand in Munich: rain shells, fleeces and hiking trousers, about 8,000 Shopify orders a month across Germany, Austria, the Netherlands and France. Returns run through Loop, support through Gorgias, reviews through Okendo and Trustpilot, and stock through an ERP that knows which factory each purchase order came from. The women's hiking trousers come back far more often than anything else, and everyone in the office has a theory.
The evidence to settle it already exists, in four systems and four languages. Loop holds a reason code and an optional comment for every returned item. Gorgias holds the conversations where customers explain what went wrong. Okendo holds reviews saying 'runs small in the waist' from people who kept the trousers anyway. The warehouse grades every return. Nobody reads all of it; the monthly report is a pie chart of reason codes, and 'too small' is the biggest slice every month.
Returns themselves are part of selling clothes online in the EU, where the 14-day right of withdrawal lets a customer send back a distance purchase without giving any reason. The expensive part is the problem hiding inside them: a waistband cut two centimeters smaller in one production batch, or a color that looks different on screen, running for months before anyone connects the comments. After a strong Black Friday, the January returns wave buries that signal completely.
Monday morning in the product channel
What the product team receives each week is a short list, not a dashboard. Below the digest are the lookups behind its first issue.
Slack, #product-quality, Monday 08:00
Returns digest · Slack
Week 36: 612 returned items, 1,904 comments, reviews and tickets read (DE, NL, FR, EN). 1. Trailline trousers, women's, sizes 36-40: 'too small in the waist'. 41 returns this week against a usual 14. 31 of the 41 shipped from purchase order PO-2607 (Porto). 'Fällt am Bund deutlich kleiner aus als meine alte Hose.' ('Much smaller in the waist than my old pair.') Suggested: measure PO-2607 samples against the size spec; add a size-guide note meanwhile. 2. Ridge shell: 'seam tape peeling after washing'. 9 returns and 3 reviews, all graded B or C by the warehouse. New this week. Suggested: supplier QA case with photos. 3. Alpaca fleece in Moss: 'color differs from the photo'. 12 of 30 returns, starting after the new product photos went live on 2 September. Suggested: compare the photo with a physical sample. Watch list: kids' rain boots, 'sole squeaks', 4 mentions. Not enough to call yet.
- themes.count(family: "TRAILLINE-W", theme: "fit.waist_small", week: 36)41 returns, 6 reviews, 4 tickets / baseline 14 a week over the last 8 weeks / second week above threshold
- erp.batch_for(order_lines: 41)31 from PO-2607 (received 12 Aug) / 6 from PO-2511 / 4 not traceable, receiving dates overlap
- warehouse.grades(returns: 41)38 graded A, resellable / consistent with fit, not a defect
- tracker.create_draft(issue: 1, owner: "product")draft ticket PROD-418 with the 41 rows, the batch split and 12 quotes / waiting for the product manager
From a return label to a line in the digest
Two models do narrow jobs here, one to classify and one to write. Every number the product team sees is computed by ordinary code from rows anyone can open.
01 Trigger · Loop, Okendo, Trustpilot, Gorgias
Returns, reviews and tickets arrive
Loop webhooks for each processed return, the review platforms' APIs, and Gorgias tickets tagged as product feedback, collected into one table per day with the order line and language.
02 Plain code · Shopify Admin API, ERP
Join each comment to variant and batch
Each comment is tied to its Shopify order line and, through the ERP, to the purchase order its stock most likely came from. Where the warehouse does not scan batches, the batch is inferred from receiving dates, and every number says how many items could not be traced.
03 AI model · Structured output
Classify against a fixed taxonomy
A small model reads each comment in its original language and assigns themes from the taxonomy below, quoting the phrase that supports each one, under a strict output schema. A comment can carry several themes, and 'no product signal' is a valid answer.
04 AI model
Surface what the taxonomy does not know yet
Comments that fit no theme are embedded with a multilingual embedding model and clustered. A cluster that keeps growing becomes a candidate theme, named by the model, which a person accepts or merges before it counts.
05 Plain code
Count, compare and apply the floor
Code counts themes per product family, variant, size, color and batch, compares each with its own baseline and with the same weeks last year, and applies the small-sample rule: below 30 units sold, the digest shows counts, never a rate.
06 Decision
Does it belong in the digest?
Thresholds the product team sets, applied by code.
- A safety word such as 'burn', 'rash' or 'snapped', in any language then the product safety owner, the same day, not the weekly digest
- A theme above its baseline for a second week then a numbered issue with a draft ticket
- A new theme, or a rise on a small base then the watch list, counted but not escalated
07 AI model
Write the digest
A model writes each issue from the computed table: theme, counts, batch split, and two or three quotes chosen for being specific, in the original language with a translation. The suggested next step comes from a short fixed list, and the model cannot state a cause.
08 Person · Slack, Jira or Linear
Product decides what changes
The product manager confirms or dismisses each issue, and decides on a size-guide note, a grading fix for the next batch, new photos or a supplier claim. Decisions go back into the sheet, so later digests show whether a fix moved the count.
The themes comments are sorted into
A fixed list keeps week-on-week counts comparable. It starts small and grows only when a person accepts a new theme.
| Theme | What it catches | Usually leads to |
|---|---|---|
| Fit: runs small or large | 'Too tight in the shoulders', 'size up', 'fällt klein aus', 'valt klein' | A size-guide note, then a grading check on the next batch |
| Fit: length and proportion | 'Sleeves too short', 'too long for my height' | Garment measurements on the product page |
| Quality: construction | 'Seam came apart', 'zip broke', 'tape peeling' | A supplier QA case with photos and batch numbers |
| Quality: material | 'Pilling after two washes', 'thinner than expected' | A material spec review with the supplier |
| Expectation: looks different | 'Color not as in the photo', 'more see-through than shown' | New photos or copy, checked against a physical sample |
| Logistics | 'Arrived damaged', 'wrong item in the parcel' | Warehouse or carrier follow-up, kept out of product counts |
| No product signal | 'Changed my mind', 'ordered two sizes', 'gift' | Nothing directly, but counted, because a rising share is a signal too |
Why raw return data misleads
These traps are why most returns reports never change a product. Each one is handled in code, not left to the model.
Reasons that say nothing
'Didn't like it' and 'other' are among the most common codes in many stores. The build does not guess a cause for them. It reads the comment if there is one, and otherwise counts the return as no signal, so vague returns cannot inflate a theme.
Codes picked to get a free return
When a faulty item comes back free and a change of mind costs a fee, some customers choose 'faulty'. The warehouse grade is the check: an item returned as faulty and graded A, with no fault found, is counted apart and never feeds a quality issue on its own. The same comparison surfaces serial returners and return-plus-chargeback double refunds, which go to a person and are never blocked automatically.
Four languages, one finding
'Fällt klein aus', 'valt klein', 'taille petit' and 'runs small' are the same finding. Classification happens in the original language, quotes stay in the original with a translation beside them, and nothing is machine-translated before it is counted.
Small numbers that look dramatic
Four returns out of nine sold is not a 44 percent problem; it is four returns. Below the floor of units sold, the digest shows counts only, and an issue needs a second week above its baseline before it is escalated.
Seasons that look like defects
January brings gift returns and the wave after Black Friday; summer brings swimwear bought in two sizes. Every comparison is against the item's own history and the same weeks last year, so a seasonal rise is not reported as a new problem.
Model, code and people, line by line
The AI model
Assign themes to each comment and quote the phrase
Short, messy, multilingual text is what a language model reads well.
Name new clusters for a person to accept
Saves someone reading two hundred comments to find the theme they share.
Write digest lines from the computed table
Prose around numbers it cannot change and quotes it did not write.
Plain code
Join comments to orders, variants and batches
A lookup that has to be exact, with untraceable items counted openly.
Count, compare with baselines, apply the floor
Arithmetic is never delegated to a language model.
Route safety words the same day
A keyword rule in every language, deliberately over-sensitive.
A person
Accept or merge new themes
The taxonomy is the product team's language, not the model's.
Decide on fixes, claims and size-guide changes
They cost money and change what customers are promised.
Is your returns app's report enough?
For a first look, often yes. Loop and AfterShip Returns both report return reasons by product, Yotpo and Okendo summarize what reviews say, and feedback analytics tools such as Chattermill or Enterpret theme free text across sources. If you want to know which products come back most and roughly why, switch those reports on and read them every week before building anything.
They stop at the edge of their own data. The returns app does not know which production batch a unit came from, the review tool never sees the warehouse grade, and neither reads the support conversation where the customer explained the problem properly. A custom build earns its place when the question is 'which batch, which supplier, which size', and answering it means joining returns, reviews, tickets, the ERP and the warehouse.
It is also a small build: a daily sync, a classification step, a weekly job and a digest, sitting next to the tools you already have. The same joined table can feed a weekly KPI brief, or tell a support agent which known issue a customer is describing, as in the order status agent.
How you would know it is working
A blueprint has no results to report, so here is what I would measure from the first week instead, on your own data.
- Issues confirmed
- Share of digest issues the product team confirms after checking. A low share means the thresholds or themes need tuning.
- Signal to decision
- Weeks between a theme first crossing its baseline and a decision in the sheet. This is the number the build exists to shorten.
- Theme count after a fix
- For each fix, the theme's returns per hundred units sold in the weeks after, against the weeks before, on bases large enough to mean something.
- Unclassified share
- Comments the taxonomy could not place. A rising share means a new theme is forming or the list has gone stale.
What a build like this costs
This is built as AI Workflow Automation, which runs $3.5K - $60K overall. A build like this one usually lands in the single-step flow tier: $3.5K - $12K, 1-2 weeks. The first working version runs on your real data well before the end of that window.
What it costs to run
Small. Classifying a few thousand short comments a week with a small model costs a few dollars, and the embeddings cost less than that. The real ongoing cost is someone reading the digest, which is the point of it.
What moves the price
- How many sources: returns only, or returns plus reviews, tickets and marketplace returns
- Whether the warehouse scans batches, or batches have to be inferred from purchase orders and receiving dates
- How many languages the comments arrive in
- Whether the output stops at a digest or also opens tickets in Jira, Linear or a supplier portal
Who this is for
- Fashion, footwear and outdoor brands where a handful of categories drive most returns
- Brands producing with more than one factory, where a fit or quality problem may belong to one production run
- Teams whose returns report is a reason-code chart that nobody acts on
- Brands selling across several EU countries, with return comments in four or five languages
Questions people ask about this
How do I find out why customers are returning products?
Read the free text, not only the reason codes, and join it to the order line. Codes say 'too small'; comments say where it is too small and compared with what. A weekly job that classifies comments, reviews and tickets per SKU, variant and batch, counts them against each item's normal level and quotes customers' own words gives the product team something it can act on.
Can AI analyze return reasons and product reviews together?
Yes, and combining them is the point: reviews catch fit problems from customers who kept the item, returns catch the ones who did not. A small model classifies each comment against a fixed taxonomy in its original language, and code does all the counting, so every number traces back to rows. The model never decides what a problem costs or what to change.
Does this work with Loop, AfterShip Returns and Shopify?
Yes. Loop and AfterShip Returns expose returns with reasons and comments through their APIs and webhooks, and Shopify holds the order lines and variants. The build reads those, the review platform and the helpdesk, and writes to a sheet and Slack. Nothing in your returns flow changes for customers or for the warehouse.
How can I reduce returns in fashion e-commerce without making returns harder?
Fix what causes them. Within the EU's 14-day withdrawal period customers do not have to justify a return, but you can stop sending them the wrong size: garment measurements on the product page, a size-guide note where an item runs small, photos with honest color, and a word with the factory when one batch is cut differently. The digest shows which of these each product needs.
How much does a returns insight system cost to build?
It is usually a compact build in the first tier of AI workflow automation: a daily sync, a classification step, a weekly job and a digest. Tracing batches through an ERP that does not record them cleanly, or adding more sources and languages, moves it up. Running costs are a few dollars a week in model usage.
Sources