Who this is for: VPs of Digital and Commerce, CTOs, and CMOs at retailers and consumer brands.
Ecommerce site search changed on 13 May 2026. Amazon retired Rufus and put an agent inside the search field itself. By 22 June, Gartner had rewritten the evaluation criteria for the entire search and product discovery category around that shift.
On 03 Aug, GSPANN also went live with its own version of new age conversational search on gspann.com.
The revenue case is real, but so is the failure record. Here’s what’s really happening:
1. Who did Amazon replace its search bar with?
Amazon built Alexa for Shopping directly into the search field, so typing a question returns an answer instead of a ranked list.
It is agentic. It can auto-buy at target prices and run scheduled repurchases. Through Buy for Me it transacts on other retailers' sites on the shopper's behalf.
The agent is free to every signed-in US customer. Prime is not required and neither is an Echo device. [1] [2]
2. Is this an Amazon experiment, or a new direction for search and discovery?
Gartner moved the goal posts on 22 June 2026. The Magic Quadrant for Search and Product Discovery now scores vendors on shopping agents, merchandising agents, product insight agents that answer questions on the product page, and MCP connectivity to outside AI platforms. Algolia, Bloomreach, Coveo, Constructor, and Netcore Unbxd were named Leaders. If you already run Adobe Experience Manager, the Algolia integration for Adobe Experience Manager search is the shortest route to that Leader tier.
PYMNTS reported in 2026 that AI shopping assistants are now the most frequently cited digital capability merchants plan to fund over the next three years. 37% of retailers named it. [3] [4]

3. Is AI shopping making people buy more?
The numbers suggest it does. Amazon reported that Rufus, its AI shopping assistant, reached 300 million users and generated $12 billion in incremental sales in 2025, and that customers who use it are 60% more likely to complete a purchase.
Walmart reports Sparky users build baskets 35% larger than non-users. Weekly active users are up more than 100% year on year and units purchased through the agent are up more than fourfold. Almost half of Walmart's app users have now engaged with it. We covered the underlying shift when e-commerce AI assistants began to outsell the search bar. [5] [6]

4. Why does it work when it works?
Because the shopper finally gets to state the constraint. LLM-style queries average 23 words against roughly four for a typed search. In 23 words the customer names a budget, an occasion, a room size, a use case.
Search was already your highest-intent surface. Across 609 million searches and $9.8 billion in revenue on 113 retail sites, searchers were 24% of visitors and 44% of revenue. They converted at 2.5 times the rate of everyone else. Conversation adds room on the one surface already carrying that number. [7] [8]

5. Can we take those Numbers to our Customers directly?
Not as stated. Both figures are company-reported, with no published methodology, sample size, or category mix.
There’s a 35% gap between people who chose to use Walmart’s Sparky and people who did not, and shoppers who engage as an assistant are already further along the conversation cycle and end up having bigger baskets.
The direction is positive and credible. [9]
6. Before we add conversation, what is the state of the search we already run?
Poor, and measurably so. 69% of shoppers go straight to the search box and 80% leave because of what comes back. Search failure accounts for 39% of site bounce. The industry zero-results rate runs 10 to 15%, so one in seven to ten queries returns nothing. 63% of e-retailers say they are dissatisfied with their own search engine. A vendor audit of 250 queries across 50 brands gave an A grade to none of them and found 85% of results pages offered no guided discovery. Search vendors have an obvious interest in that last result. It is also consistent with every other number here. Our white paper on e-commerce site search performance walks through the same diagnostic on a real catalog. [10] [11]

7. What changes when we put a conversational layer over that?
The failure mode changes, and it stops being visible.
Every query ends in one of three states. Answered, where the system finds what exists. Refused, where it returns nothing and says so and the most dangerous: Fabricated, where it produces a fluent, confident recommendation the catalog cannot support.
This is where Keyword search won. It could never produce a fabricated result. A conversational agent over the same thin catalog produces all three, and therefore requires specific guardrails, training, and governance. [12]

8. Why is a fabricated result worse than refused?
Because you can measure a refusal. Zero-results is a number on a dashboard. Someone owns it and it appears in a weekly review.
A confident wrong answer produces no error and no alert. Nothing in the log reads as failure. The shopper gets a recommendation that does not fit, does not buy or buys and returns, and your instrumentation stays quiet.
Conversation over a thin catalog removes the search box's error message. The failure rate stays where it was. The reporting is what changes. Every system that lets an agent speak for the business runs into the same question of whether AI agents can be trusted with enterprise data. [13]
9. Has this blown up on anyone, or is that just caution?
It has blown up on most people who tried it.
Stacker reported in 2026 that 74% of organizations have been forced to shut down or roll back a live AI customer communications agent. The leading consequences were increased support queue volume (35%) and brand reputational damage (34%). Of those already in production, 75% have had at least one governance rollback. The same rollback rate drove our analysis of AI agent governance in customer experience.
Klarna, for example, went AI-first on service alongside a roughly 40% headcount reduction, then rehired humans when the system underperformed. [14]

10. If it says something wrong about a product, are we legally on the hook?
Yes. On 12 May 2026, the Higher Regional Court of Hamm held a company liable for misleading statements made by its own AI chatbot, which had described its practitioners using professional titles that did not exist.
The court treated the chatbot's output as a misleading commercial act by the company itself under German unfair competition law. [15] [16]
11. What about other companies’ agents shopping on our site?
That fight is already in federal court.
Amazon sued Perplexity over its Comet shopping agent making purchases on Amazon. The complaint alleges Perplexity disguised automated agents as human users and kept accessing customer accounts after being told to stop.
Whatever the outcome, you will soon need to know which of your sessions are people and which are agents, and today most of us cannot tell. [17] [18]
12. So what are our options?
There are five paths to conversational search.
- Amazon's packaged product gives you the fastest credible path, at the cost of strategic dependency on a competitor and no public pricing.
- Your search vendor's agentic layer avoids a replatform but makes you dependent on their roadmap.
- Your commerce platform's native tools are often already paid for and capped at the platform's ceiling.
- Building it gives full control on a long timeline.
- Fixing deterministic search first costs the least and shows a result this quarter. It is also the precondition for the other four. [19]

13. What do we get and give up if we take Amazon’s version?
Speed, in exchange for two real concessions.
AWS packaged the Alexa for Shopping technology as the Agentic Shopping Assistant, built on Bedrock, AgentCore, and OpenSearch, with a launch window of about sixty days.
The reference customer is Kate Spade, whose AI Gift Concierge went live on Bedrock AgentCore before the packaged product existed. That tells you the components are proven and the package is new.
AWS states retailers keep control of their customer data, catalog, and business rules. The concessions: pricing is private offer only with no public list price, and the company supplying your discovery layer runs an assistant whose Buy for Me feature shops your competitors. [20] [21]

14. Can we just switch on what we already pay for?
Often yes, and this is the most underused option on the list.
Salesforce ships Agentic Commerce Search natively in B2C Commerce. It builds a commerce-tuned small language model on your own catalog and shopper data, then exposes it through headless APIs to retailers running Shopify, commercetools, SAP Hybris, and Adobe Commerce. Shopify ships every AI feature free on every plan. We looked at which Salesforce Agentforce commerce agents are safe to put in front of customers first.
Adobe offers LLM-powered discovery and semantic search, and most Adobe Commerce customers have never activated the AI features already inside their license. That gap is a switch nobody flipped, not a budget problem. [22] [23]
15. Should we build it the way Walmart did?
Only if you are honest about which Walmart you are.
Sparky is a retrieval-augmented system. Walmart consolidated it from 200 bots down to four using MCP, deployed it across outside language models through open protocols, and deliberately pulled back from pure OpenAI dependence.
Walmart optimizes it for what it calls answerability: how well structured product data matches the intent of a natural-language question. That word is the whole lesson. Walmart made the catalog answerable first. The model came second. That is the same sequence behind automated product data management with Salsify PIM, where attribute completeness set the pace for everything downstream. [24] [25]

16. What should we measure before we sign anything?
Two numbers, and you can have both in a fortnight.
First, your zero-results rate broken out by category. That tells you how much of your catalog is already unreachable.
Second, take the fifty questions your customers actually ask and count how many can be answered from structured attributes rather than prose.
That second number is your ceiling under every option on the table. No vendor demo will tell you what it is, because the demo runs on their catalog. The wider discipline behind it is covered in our white paper on data accuracy in AI-driven systems. [26]
GSPANN's Take
- The category is being sold as an interface upgrade. It is a retrieval change, and retrieval is only as good as the attributes underneath it.
- The failures trace to three things: thin catalogs, absent governance, and support teams cut before the system could carry the load. Conversation is rarely the broken part.
- The mechanism nobody prices in: refusals are instrumented and fabrications are not, so the dashboard gets quieter while the catalog stays thin.
- Walmart named the actual variable when it optimized Sparky for answerability instead of for the model. Answerability is where our own product data work starts, and we saw it again when we put conversational search on gspann.com: attribute coverage set the ceiling on what could honestly be answered.
- Measure your zero-results rate before you take a demo. It is the only number in this decision that is already true.
Nobody sends a report when the assistant invents a product that fits. The shopper asks for something specific, gets a fluent recommendation that may be incorrect or completely made up. We recommend you proceed with the new conversational search but with better and clean data, strict governance, and caution.
All References
Ref 2: https://www.retaildive.com/news/amazon-ai-alexa-for-shopping/820218/
Ref 3: https://constructor.com/blog/gartner-magic-quadrant-2026-leader
Ref 4: https://www.pymnts.com/news/artificial-intelligence/2026/ai-shopping-assistants-top-retail-budgets/
Ref 5: https://ppc.land/amazons-ai-shopping-assistant-drove-12-billion-in-sales-for-2025/
Ref 6: https://www.digitalcommerce360.com/2026/05/22/walmart-sparky-agent-ai-sales-supply-chain/
Ref 7: https://www.localogy.com/2026/02/soci-study-llm-queries-are-6x-longer-than-search-queries/
Ref 9: https://www.modernretail.co/technology/walmart-says-ai-users-build-35-bigger-baskets-than-others/
Ref 10: https://www.nosto.com/blog/new-search-research/
Ref 11: https://zoovu.com/resources/ecommerce-search-report
Ref 12: https://zoovu.com/resources/ecommerce-search-report
Ref 13: https://www.nosto.com/blog/new-search-research/
Ref 15: https://www.lexology.com/library/detail.aspx?g=803b6132-e0c7-4971-aeac-c0f423137e0a
Ref 19: https://www.aboutamazon.com/news/aws/aws-agentic-shopping-assistant-retailers
Ref 20: https://www.aboutamazon.com/news/aws/aws-agentic-shopping-assistant-retailers
Ref 22: https://www.salesforce.com/blog/new-agentic-commerce-search-capabilities/
Ref 23: https://business.adobe.com/products/commerce/ai-commerce.html
Ref 24: https://stellagent.ai/insights/walmart-sparky-super-agent
Ref 25: https://constructor.com/blog/gartner-magic-quadrant-2026-leader






