← BACK TO THE BLOG

AI for ecommerce: where it moves the needle and where it's just noise

Laptop showing the words online shopping, with a miniature shopping trolley holding a phone and several paper bags and cardboard boxes on the keyboard

The conversation about AI for ecommerce usually starts in the wrong place: mass-generating product descriptions, or building an assistant that chats like an over-eager shop floor clerk. Meanwhile the store's internal search returns zero results on one in five queries and nobody is looking. This guide does the opposite: it covers the four levers where AI changes real numbers in an online store — search, recommendations, returns and customer service — what data each one demands, and the order in which to tackle them.

A caveat up front, for honesty's sake: almost none of these levers are "AI" in the spectacular sense of the word. They are unremarkable models dropped into a flow that already exists. That is precisely what makes them pay for themselves.

What does AI actually bring to an ecommerce business?

An online store has an advantage almost no other business has: everything is measured and everything is already digital. Every search, click, abandoned cart and return leaves a record. That makes ecommerce the ground where AI shows results fastest, because the data already exists and the effect shows up in days, not quarters.

The flip side is that margin rules. In a store running a 6% net margin, a model that gets 3% more recommendations right can be irrelevant if it costs €2,000 a month to maintain. The right question isn't "does this work?" but "does this move enough euros to pay for its own upkeep?".

Apply that filter and four areas survive almost every time:

| Lever | What it improves | Data it demands | Sign that it's working | |---|---|---|---| | Internal search | Conversion of people already intending to buy | Search history and a clean catalogue | Zero-result searches drop | | Recommendations | Average order value and units per order | Order history with customer and product | Module click rate and basket size rise | | Returns | Margin leaking out of the back door | Return reason per order line | Returns for size or wrong description drop | | Customer service | Cost per contact and response time | Ticket history and real order status | Fewer contacts escalate to a human |

And one that rarely survives its first year: bulk generation of product content with no review. Not because the copy comes out badly, but because multiplying mediocre product pages doesn't sell more and does pollute the catalogue — the very asset the other three levers depend on.

Why is internal search the first lever?

Because anyone using a store's search box has already decided to buy. It's the hottest traffic you have and, in most mid-sized stores, the worst served. The usual pattern: search does literal text matching, so "kids rain boots" never finds the "waterproof children's wellingtons" you have in stock, and that shopper leaves.

What semantic search changes — models that represent product and query by meaning rather than by letters — is concrete:

The prep work is unglamorous and accounts for 70% of the outcome: normalising catalogue attributes (size, colour, material, compatibility), deciding what to do with out-of-stock items, and clearing duplicates. A semantic model on a dirty catalogue returns dirty results, faster.

Before you buy anything, export the 200 most frequent searches of the last three months and count how many return zero results. That number, not a demo, is your business case.

Do product recommendations actually work?

Yes, but less than the industry promises and very unevenly depending on the catalogue. Three kinds of recommender, from least to most demanding:

1. Rules and co-occurrence. "People who bought this also bought that", computed over order history. It's a SQL query, not AI, and on small catalogues it performs nearly as well as a model. Always start here. 2. Collaborative filtering. Learns patterns across customers who behave alike. It needs volume: with few repeat customers, the model has nothing to learn from. 3. Content and context models. Combine product attributes, session history and timing. They justify their cost on large catalogues with high turnover.

The classic trap: recommending what the customer was going to buy anyway. It lifts the dashboard number and adds nothing. The only honest test is an A/B test with the module switched off for a control group, measuring revenue per session — not clicks on the module — across at least two full purchase cycles.

Watch out for the narrow-catalogue effect too: a recommender optimised for the short term concentrates sales on the same fifty products and kills the long tail. If your margin lives in the tail, measure that explicitly.

What can AI do about returns?

This is the forgotten money. In fashion, online return rates sit in ranges that wreck margin; in electronics they're lower but each case costs far more. And hardly any store analyses its returns beyond counting them.

Three uses that hold up:

Returns are also an input to demand forecasting: stock that comes back, and when, changes the picture of real availability, and the rest of the chain benefits — as we cover in AI in logistics.

When does a customer service assistant make sense?

When the volume of repetitive contacts is high and the answer lives in a system the assistant can query. That second half is everything. In ecommerce, between 50% and 70% of contacts are variations on four questions: where is my order, how do I return it, when do I get my refund, and is this item available.

An assistant wired into the order system and the carrier answers all four with the exact data. An assistant that has only read the FAQ answers in generalities and achieves something worse than not existing: the customer writes twice. The difference isn't the model, it's the integration.

Three rules we always apply:

The detail on use cases, limits and real cost is in chatbots for business and, if your channel is WhatsApp, in WhatsApp chatbots.

Where do you start on a realistic budget?

An order that works for stores turning over between one and twenty million:

1. Weeks 1-2: measure what already happens. Zero-result searches, return rate per SKU, contact volume by type, average order value with and without the recommendation module. Without this baseline there's no way to know later whether anything worked. 2. Weeks 3-6: fix the catalogue. Normalised attributes, rewritten pages for the most-returned products, size charts reviewed. It's the dullest task and the one with the most direct return. 3. Weeks 6-10: one lever, with a control. The one your measurement points to, not the one that's fashionable. With a control group and a business metric agreed in advance. 4. After that: the next one. Only if the first held its result for two consecutive months and someone on your team can run it without the vendor.

A word on cost: the infrastructure behind these levers is cheap; what costs money is catalogue upkeep and periodic review. Budget people's hours, not just licences. How that compares with other projects is in what an AI project costs.

Frequently asked questions

Do I need a lot of traffic for AI to pay off in my store?

For learning-based recommendations, yes: without thousands of historical orders there's no pattern to learn. For semantic search and return classification, no — they work on modest catalogues and volumes because they lean on content rather than aggregate behaviour.

Can I just use the AI my ecommerce platform includes?

It's usually the sensible starting point: zero marginal cost and zero integration. The limit appears when you need your own logic — your stock rules, your margin criteria, your catalogue with unusual attributes — or when you want to take your data with you if you switch platform. At that point the logic should live outside the product.

What about my customers' data?

It remains personal data even when a model processes it. Before sending purchase histories or conversations to an external provider you need to review legal basis, processor agreements and where the data sits. The techniques for working without exposing customers are in data anonymisation.

How long before the effect shows?

For search and customer service, weeks: traffic is continuous and the effect appears in the following week's metric. For returns and recommendations, one to three months, because they depend on the purchase cycle and the returns window. Be sceptical of anyone promising measurable results in days.

If you recognise your store in two or three of these symptoms — search that doesn't find, returns nobody analyses, support drowning in "where is my order" — the next step isn't buying a tool: it's measuring what each one costs you in euros. That's what we do in the audit, with your systems and your numbers on the table; and if you'd rather test the idea in a short conversation first, let's talk.

Shall we apply it to your case?

The 360° AI Audit turns these ideas into a concrete plan for your company: three weeks, fixed price and the full picture of your AI before spending a euro.

See the 360° Audit Let's talk