AI used to be wrong inside a chat box.
Now it can be wrong while buying something, sending a message, booking your day or acting inside an account.
That changes the cost of a mistake.
Recent examples documented by The Wall Street Journal show just how quickly convenience can turn into exposure. Personal AI agents have shared sensitive information in the wrong place, made purchases without clear confirmation and confused everyday scheduling tasks. One agent accidentally posted a startup CEO’s bank-balance information into a company Slack channel. These are individual incidents, not an industry-wide failure rate, but they reveal a problem that becomes harder to ignore as AI moves from answering questions to taking actions.
Shopping is where that problem becomes especially interesting because the agent is no longer merely summarizing information. It has to decide what matters, compare imperfect sources, infer what the user wants and eventually influence where money goes.
Researchers at Wharton Generative AI Labs tested that process roughly 26,000 times. When an AI shopping agent saw only a product page, its recommendation was relatively stable. Add a review, change the order of sources, inject a piece of user memory or alter how information is retrieved, and the final choice became less predictable. The more real-world context entered the decision, the less consistent the recommendation became.
That is an uncomfortable result for the next phase of commerce.
AI platforms are moving in exactly the opposite direction: toward deeper access and greater delegation. OpenAI’s Agentic Commerce Protocol already connects ChatGPT with structured product data from retailers including Target, Best Buy, Lowe’s, Nordstrom, Sephora, The Home Depot and Wayfair, while Shopify catalog data can feed product discovery directly into ChatGPT. Walmart has also launched an in-ChatGPT experience with account linking and payments.
The infrastructure is being built for AI to sit much closer to the transaction.
But retailers are not uniformly comfortable giving external agents that power. Amazon has blocked Meta’s Muse from browsing and buying on its retail site, highlighting a new question for the web: when is an AI agent a legitimate customer representative, and when is it just another bot asking for access?
Consumers appear to be drawing their own boundary. Product.ai surveyed 1,463 U.S. online shoppers in 2026 and found that 43% had used AI for product research in the previous 90 days. Among the 623 people who had used AI that way, 86% checked the recommendation against another source before buying. Only 14% said they trusted the AI recommendation without verification.
That gap between using AI and trusting AI enough to act without checking may be the real constraint on agentic commerce.
The next generation of shopping agents will therefore need more than better models. They need clear permission boundaries, reliable product data, traceable sources, explicit confirmation for consequential actions and an audit trail showing exactly what the agent did. OpenAI’s own earlier commerce design reflected that logic by requiring users to explicitly confirm purchase steps and limiting payment authorization to a specific amount and merchant.
The important benchmark is shifting.
It is no longer just whether an AI agent can complete the task.
It is whether users can trust it to know when to act, when to ask and when not to click at all.