Insights
Twelve generative AI features shipped in commerce products, what each does, who checks the output and where it has gone wrong in public. Generative AI has shipped in commerce products in three places: drafting content a merchant publishes, helping shoppers find and judge products, and speaking or acting for the busines
Generative AI has shipped in commerce products in three places: drafting content a merchant publishes, helping shoppers find and judge products, and speaking or acting for the business in chat and checkout. The first two hold up better, though both have public failures. What separates the durable uses is not whether a check exists on paper, but whether the output sits next to something a person can compare it with. The third is where generated words become the business's commitment on a price, a policy or a purchase, and the designs that shipped there keep that commitment in structured systems with a person confirming it.
Last reviewed September 28, 2026.
We included only features documented in the company's own announcement, help center or credible reporting that we opened, and we date each source. Vendor figures are labeled as the vendor's. Where a source describes a chatbot without saying it is generative, we say so.
Who catches the mistake decides where generative AI belongs
Every generative feature produces some wrong outputs. The useful product question is who sees a wrong output first, what it costs by then, and what the product gives that person to catch it. The table sorts twelve shipped uses by that question.
| # | Use | Shipped in | Who sees the output first | The expensive wrong output | The check the product ships |
|---|---|---|---|---|---|
| 1 | Product copy drafts | Shopify Magic, Amazon listing tools | Merchant | An invented spec or benefit published as the retailer's claim | Merchant edits before publishing |
| 2 | Listings from a photo | eBay magical listing | Seller | Wrong item details or category | Seller edits the generated listing |
| 3 | Catalog data at scale | Walmart | Not publicly described | A wrong attribute that search and operations rely on | Not publicly described |
| 4 | Admin assistant that proposes changes | Shopify Sidekick | Merchant | A wrong edit to products or orders | Changes shown for review before they apply |
| 5 | Review summaries | Amazon review highlights | Shopper | A minority complaint presented as consensus | Attribute buttons that open matching reviews |
| 6 | Shopping assistant answers | Amazon Rufus, now Alexa for Shopping; Zalando Assistant | Shopper | A wrong product fact, or a sponsored result read as advice | Results are real product pages |
| 7 | Generated gift guides | Etsy Gift Mode | Shopper | Irrelevant suggestions | Every suggestion is a real listing |
| 8 | Virtual try-on | Google Shopping | Shopper | A shopper reading fit or size from the image | Google states the image does not indicate fit or size |
| 9 | Generated recipes and food images | Instacart | Shopper | An impossible recipe presented as the retailer's content | None visible in the reporting |
| 10 | Replies to shoppers | Shopify Inbox | Merchant for suggested replies; shopper for the Inbox agent | A wrong answer on shipping, returns, sizing or availability | Merchant edits suggestions; agent hands off to staff |
| 11 | Customer service on refunds and disputes | Klarna AI assistant | Customer | A refund or dispute handled badly | A route to a human, which Klarna reinforced in 2025 |
| 12 | Buying on the shopper's behalf | Amazon auto-buy, Google agentic checkout, ChatGPT Instant Checkout | Shopper, at confirmation or afterwards | An unwanted purchase at the wrong price | Shopper-set target, confirmation step or cancellation window |
Uses a merchant checks before shoppers see them
These keep a person between the model and the customer. They are the safest category on paper, and the evidence shows why "on paper" matters.
1. Product copy drafts: Shopify Magic and Amazon's listing tools
Shopify's product-description generator turns a title and a few features or keywords into a suggested description in the admin. Shopify's help center (accessed September 2026) says the merchant is responsible for the accuracy of what they publish, and warns that generated text can include benefits the merchant never listed and facts drawn from content about similar products. Amazon launched a comparable tool for sellers in September 2023: a short description in, a title and description out, for the seller to refine or submit directly. Amazon said that in testing, the majority of sellers were using the generated content directly.
Where it goes wrong: the dangerous invention is the plausible one. A made-up material or compatibility claim reads like any other line of copy, and once published it becomes the retailer's statement. Review also gets skipped entirely. In January 2024, Futurism reported Amazon listings whose titles were ChatGPT refusal messages, which suggested sellers were using ChatGPT to write them; Amazon removed the listings.
What to take from it: a review step is only a check if the interface makes the risky parts visible. Showing which statements were not in the merchant's input is a different product from a text box pre-filled with fluent copy.
2. Listings from a photo: eBay magical listing
eBay announced in September 2023 a tool that takes a photo in the app and generates a title, description and item details such as category and release date, and can pair with other eBay tools to suggest a price and shipping cost. Sellers could edit what it produced.
Where it goes wrong: resale listings carry condition, edition and authenticity details a photo may not establish. The generated draft can be confident about precisely the details a buyer will dispute.
What to take from it: the seller holds facts the model cannot see. The draft should ask for them rather than guess them.
3. Catalog data at scale: Walmart
On its August 2024 earnings call, Walmart said it had used large language models to create or improve more than 850 million pieces of product catalog data, as CIO Dive reported. Executives said catalog data quality affects how customers find and buy products, how inventory is stored and how orders are delivered.
Where it goes wrong: we found no public description of how Walmart checks this data, so we cannot say how it performs. The structural point stands without that detail. An attribute error here is not seen as generated text; it can surface later as a wrong filter result or a wrong decision about storing or delivering the item.
What to take from it: generative work behind the interface still needs an owner and a sampling check, because its errors surface far from where they were made.
4. An admin assistant that proposes changes: Shopify Sidekick
Sidekick, Shopify's assistant for merchants, can analyze store data, generate content, edit products and help manage orders. Its help center page (accessed September 2026) says it presents changes for review before applying them, and that longer tasks run in the background and notify the merchant when ready for review.
Where it goes wrong: a review screen for a bulk change is only as good as the merchant's ability to see what changed. We have not seen the interface for large edits, so we cannot judge it.
What to take from it: proposing rather than applying is the right default for an assistant with write access. The design question is how the proposal shows its scope.
Uses shoppers can check against evidence
These address the shopper directly, but the output points at something real: reviews, listings or product pages.
5. Review summaries: Amazon review highlights
Amazon introduced AI-generated review highlights in August 2023: a short paragraph on the product page summarizing features and sentiment that recur across reviews, drawn only from verified purchases, with attribute buttons that surface reviews mentioning a topic such as ease of use.
Where it goes wrong: in December 2023 Bloomberg reported sellers' complaints that summaries overweighted rare negatives. One example, summarized by Slashdot, was a board game rated 4.7 stars where a handful of ease-of-use complaints took up about a third of the summary. The link back to reviews does not correct this, because a shopper who trusts the summary does not open them. The party harmed, the seller, has no check at all.
What to take from it: a summary that links to its evidence is checkable, not checked. Proportion is part of accuracy, and it is the part a summary most easily loses.
6. Shopping assistant answers: Amazon Rufus, now Alexa for Shopping, and Zalando Assistant
Amazon's Rufus answered product questions and recommended items inside the shopping app from 2024. Tests found errors that matter for commerce: TechCrunch (March 2024) got a women's vest when asking for men's leather jackets and found Rufus could not check an order or start a return; Marketplace Pulse (November 2024) found answers to "cheapest" that were not the cheapest. Amazon also began testing sponsored ads in Rufus in September 2024, noting it may generate text to accompany existing ad copy. In May 2026 Amazon renamed Rufus and merged it into Alexa for Shopping, which puts AI-generated overviews at the top of search results in the app.
Zalando took a narrower brief. Its assistant, in beta across all 25 markets from October 2024, takes questions such as what to wear to an occasion in a given city and season, and answers with items from the assortment.
Where it goes wrong: a product the shopper can open is self-correcting in a way a stated fact is not. A wrong "cheapest" or a mis-stated spec looks exactly like a right one. When the assistant can also carry advertising, the shopper cannot tell from the answer whose objective produced it.
What to take from it: styling advice that returns products leaves the shopper to judge. Factual claims about price or specification need to come from the catalog, visibly, and a sponsored item needs to be marked where the model mentions it.
7. Generated gift guides: Etsy Gift Mode
Etsy launched Gift Mode in January 2024: a short quiz about the recipient, occasion and interests, returning gift guides built around more than 200 personas. Etsy said it combined GPT-4 with its existing machine learning and human curation.
Where it goes wrong: we found no public reporting of failures. The design explains why a failure would be cheap: the generated layer organizes the catalog into personas, and everything the shopper can buy is a real listing they can inspect.
What to take from it: using generation to structure browsing, rather than to state facts, keeps the model away from the claims a retailer has to stand behind.
8. Virtual try-on: Google Shopping
Google extended try-on to shoppers' own photos in May 2025, generating an image of a garment on a full-length photo the shopper uploads. Its help page (accessed September 2026) says generated images can contain mistakes, that the image does not determine or guarantee fit, that size availability is set by the merchant, and that shoppers should use size charts, reviews and product details to choose a size.
Where it goes wrong: fit is a central question when buying clothes, and an image of the garment on your own body invites a fit judgment the feature disclaims. The disclaimer moves the check to the shopper.
What to take from it: Google scoped the feature honestly. A retailer building try-on has to decide whether a disclaimer is enough at the point where the shopper chooses a size, or whether size guidance belongs beside the image.
9. Generated recipes and food images: Instacart
Instacart has shown shoppers AI-generated recipes. In February 2024, 404 Media reported Instacart recipes with measurements and ingredients that do not appear to exist, illustrated with AI-generated images that were not disclosed as such. We have not checked what Instacart shows today.
Where it goes wrong: content in the retailer's voice, with no marker that it was generated and no evidence behind it, carries the retailer's credibility. The shopper cannot compare it with anything.
What to take from it: this is the counter-example for the category. Inspiration content feels low-stakes, but undisclosed generation removes the shopper's only reason to be skeptical.
Uses that speak or act for the business
Here the output is a statement or an action the business is answerable for: a policy, a price, a refund, a purchase.
10. Replies to shoppers: Shopify Inbox, suggested and sent
Shopify Inbox shows both ends of the authority scale in one product. Suggested replies are drafted from store information such as product listings, policies and shipping settings; staff are told to read and edit them before sending. The Inbox agent (help center, accessed September 2026) is opt-in and chats with customers directly: it answers product questions, recommends, looks up orders and adds items to the cart, and hands off when a customer asks for a person or a conversation needs staff. Shopify's guidance says the merchant is responsible for the accuracy of what the agent tells customers.
Where it goes wrong: the same sources feed both modes, but only one has a person reading before the shopper does. The public failure of this category is older. In December 2023 a car dealership's ChatGPT-based website chatbot, built by Fullpath, was talked into agreeing to sell a new SUV for one dollar and calling it a binding offer. Fullpath's own account describes a deliberate manipulation, and the AI Incident Database entry, which collects the reporting, records that no sale took place.
What to take from it: when a merchant switches from suggested to automatic, what changes is who checks, not the model. Price, availability and policy answers should be read from the systems that own them, and anything resembling an offer should be out of the agent's scope.
11. Customer service on refunds and disputes: Klarna
Klarna launched its OpenAI-powered assistant in February 2024 to handle refunds, returns, payment issues, cancellations, disputes and invoice inaccuracies, with live agents still available. Klarna claimed it handled two-thirds of service chats in its first month. In May 2025 its chief executive said cost had been too dominant a factor and the result was lower quality, as CX Dive reported, and Klarna began recruiting human agents again while keeping the assistant.
Where it goes wrong: Klarna has not published which kinds of conversation fell short, so we cannot say whether the problem was accuracy, resolution or tone. What it chose to change is public: a reliable route to a person.
The legal backdrop matters here. In February 2024 a British Columbia tribunal held Air Canada liable for a chatbot's wrong answer about bereavement fares, rejecting the argument that the chatbot was responsible for its own actions and finding that customers should not have to double-check one part of a website against another (McCarthy Tétrault summary). The decision does not say whether that chatbot was generative, and it is one Canadian small-claims ruling, but its reasoning applies to any automated answer on a commercial site.
What to take from it: in service, the model is explaining commitments the business has already made. Each answer about a refund or policy is the business's answer.
12. Buying on the shopper's behalf: Amazon, Google and ChatGPT
Three shipped designs place the check differently.
- Amazon added auto-buy to Rufus in November 2025: the shopper names a target price, Amazon buys with the default payment method and address when the price is reached, sends a notification and allows free cancellation for 24 hours. Alexa for Shopping carries on price-triggered buying and adds scheduled restocking, as GeekWire reported in May 2026.
- Google's agentic checkout, rolling out from November 2025 with merchants including Wayfair and Chewy, lets a shopper set size, color and budget on a tracked item, and says it asks permission and buys only after the shopper confirms purchase and shipping details.
- OpenAI launched Instant Checkout in ChatGPT in September 2025 on a protocol built with Stripe, under which merchants can accept or decline orders and handle fulfillment and returns. By March 2026 native checkout inside ChatGPT had been withdrawn in favor of merchant apps and referrals, according to Forrester, and Walmart said in-chat purchases had converted at about a third of the rate of shoppers sent to its own site, as MarTech reported. Walmart moved to running its own assistant, Sparky, inside ChatGPT with its own checkout.
Where it goes wrong: we found no public reports of these systems buying the wrong item. The ChatGPT retreat is evidence about something else: when checkout moved away from the retailer, a large retailer chose to take it back. That is a conversion figure Walmart reported, not a finding about model error.
What to take from it: the delegated purchases that shipped keep the model away from the commitment itself. The shopper sets the terms in structured fields, the merchant's systems accept the order, and a confirmation or cancellation window sits around the moment money moves.
What the evidence changes about where AI belongs
We started with a hunch: uses hold up when a wrong output is cheap to catch, and go wrong when the model speaks for the retailer at a moment of consequence without a check. The research supports the second half and qualifies the first in three ways.
Cheap to catch is not the same as caught. Merchant review is the check in uses 1 to 4, and the vendors' own signals suggest it can be light: Amazon said most test sellers used generated listings directly, and sellers apparently using ChatGPT let refusal messages reach live titles. The fix is in the interface, not the policy: show what the model added, and make the risky fields hard to accept unread.
Evidence nearby helps only if it is proportionate and marked. Review summaries, try-on and assistant answers all sit next to real evidence, yet each has a known failure the evidence does not repair: minority complaints as consensus, fit inferred from an image, sponsored placement inside advice. The shopper can check, but the interface gives little reason to.
At moments of consequence, the designs that shipped keep the commitment in structured systems. Auto-buy, agentic checkout and Sidekick all turn the model's work into a proposal that structured systems execute, with a person setting or confirming the terms. The failures, from the dealership chatbot to the undisclosed recipes, are cases where generated words became the business's statement with nothing between them and the customer.
This matches the principle on our commerce and transactional products page: AI can support discovery, configuration and service, but recommendation objectives, evidence, inventory truth and commitment authority must remain explicit.
Questions to answer before you ship a generative feature
| Question | Why it matters |
|---|---|
| Who sees the output before a customer relies on it? | If the answer is nobody, the feature is speaking for you. |
| What is the most expensive wrong output, and when is it discovered? | A wrong adjective costs little; a wrong return window is found after purchase. |
| Which system holds the truth for price, stock, policy and order state? | The model should read these facts and show them, not restate them in its own words. |
| What can the reviewer or shopper compare the output with? | A link to evidence helps only if the proportion and source are visible. |
| Is generated content marked, and are sponsored items marked where the model mentions them? | Undisclosed generation removes the reason to be skeptical. |
| What does the person confirm, and how do they undo it? | Delegated actions need structured terms, confirmation and a recovery path. |
| Who owns quality after launch? | Klarna's own judgment of service quality led it to reinvest in human support. |
Where this work sits
Deciding which of these uses belongs in a product, and where its check goes, is a product decision before it is a model decision. Tcules works on commerce and transactional products, where the purchase has to stay coherent from discovery through fulfillment and recovery, and on AI Product UX, where authority, evidence, review and recovery are designed into the interface.
The commerce product-experience guide works through the evidence a shopper needs at each stage of a purchase. For examples outside commerce of products that place a human check at the moment of consequence, see AI trust, control and human review examples. If your team has not yet agreed what an AI feature may do or who owns its output, an AI Product UX Readiness Assessment is a better starting point than interface design.