Search, scroll, filter, compare, add to cart — for thirty years online shopping was a sequence of clicks. The next shopper simply says what they need, and an AI assistant does the rest. That shift is agentic commerce. This site is the first human-verified look at how well it already works for Indian shoppers.
Every step of online shopping was a click: search, scroll, filter, compare, cart. The next shopper says what they need — “a sunscreen for oily skin under ₹1,000” — and the assistant does the rest. The interface itself is changing.
For the first time in the history of commerce, buying will happen across five foundational surfaces at once — the Big 5 AI labs — and possibly more. No single storefront and no single search box owns the shopper any longer.
E-commerce ended the manipulation of the shop-floor salesman by letting us do our own research. Then the system caught up: paid reviews, sponsored creators and generated content now shape what you think — without you knowing it.
We believe agentic commerce solves this. A highly intelligent entity that knows your needs, and is not moved just because your favourite content creator posted about a product or because a brand paid for a review you happened to see.
A language model’s natural input and output is words. For at least some product verticals, the return on visual advertising will keep falling — the shopper those ads were made for is no longer looking at them.
Commerce has never been contested by five foundational model makers at once. Each arrives with a different history, a different relationship with advertising — and a different reason to win the shopper.
The newest name in commerce and advertising — and the maker of the strongest models. It has never sold a shopping ad; it is learning the shelf in public, and fast.
The tried and trusted player: search, shopping and ads have been Google’s for two decades. For the first time in its existence it faces competition on its own ground — and it answers on two surfaces.
A new entrant whose appetite for the consumer market is not yet clear. Yet many of us already turn to Claude for buying decisions — so it is on the shelf, whether it planned to be or not.
Compute, talent density and data at enormous scale, plus years of advertising experience through X. Its latest models trail the Big 3 closely — and it arrived with a business model already attached.
The largest advertising machine ever built, now shipping models that closely trail the Big 3 — with billions of shoppers already inside its apps and years of ad data behind every answer.
An average shopper doesn’t ask one perfect question. They ask, get filtered, weigh the options, hunt for a deal and end up on Amazon.in or Flipkart. We followed every answer through those six steps — then opened everything it pointed to.
Did the assistant fully understand what the shopper wanted — did it ask clarifying questions before it recommended anything, or did it guess?
Did it honour every filter in the question? “Sunscreen for oily skin within ₹1,000” carries two — skin type and price — and we checked both.
Did it explain the pros and cons of the shortlisted products, or simply name three and stop?
Did it fetch the best offers, and when we opened the merchant page, were those offers actually there?
Did it hand over Amazon.in and Flipkart links — and did the prices and reviews it quoted from those two surfaces match what the pages showed?
For every other merchant or brand link it gave, did the page open on the right product, and did the price and reviews match the brand or merchant site?
Nothing here is synthetic or self-graded. Every figure on this site traces back to a human asking a question, opening a page and writing down what they saw — in India, for Indian shoppers.
Across 3,112 base answers and 22,894 follow-up probes, the basics of agentic shopping are already solid: constraints respected, budgets honoured, prices quoted, buy links handed over. Each ring is a share of all answers.
links from the best assistant opened on a live product page — ChatGPT's 91%. The shelf is reachable when the model is careful.
Source scores for Gemini, Claude, AI Overviews, Grok and Meta AI: the pages they cite exist and resolve. Citation hygiene is largely solved.
of price misses were within 10% of the shelf price; fewer than 3% were off by more than half. When AI gets the price wrong, it is usually a little wrong.
16,497 merchant pages, opened by hand and compared to what the assistant claimed. Each dot is one percent of them.
of quoted prices matched the merchant page — 6,099 of 8,309 checked. ChatGPT reached 86%; Grok 79%. When they missed, the median gap was 8%.
of quoted ratings and review counts matched — 3,917 of 6,412. Claude led at 87%; AI Overviews trailed at 29%. Ratings are the claim to double-check.
Amazon.in and Flipkart price quotes that matched. Marketplace prices move daily — the one number to re-check before checkout.
Reliability is not one number — it depends on what you're buying and who you ask. Skincare and supplements are already safe ground; electronics is where the shelf moves fastest.
The most reliable pairing in the study — 455 pages, with 94% of ratings true as well.
Claude's supplement answers are precise: 92% of pages live, 85% of prices matched.
Grok quotes chair prices almost perfectly (36 checks); Gemini's chair ratings were 85% true.
The hardest aisle: SKUs churn weekly. Even here ChatGPT held 88% — the gap is closable.
The products each assistant named most often for the same question — and how often it reached for them. Different assistants, genuinely different taste.
Counts are how many of that assistant's answers to this question named the product. Claude and Grok were captured on fewer runs in Wave 1. Auto-advances every few seconds; hover to pause.
{{ dimDesc }}
Composite, Truth, Buy path, Filter and Sources come from the metrics pipeline (codebook v0.24, Wave 1 India). Everything else is computed straight from the India capture database — 3,112 answers, 16,497 merchant pages, 4,340 deal checks. † thin sample — Grok 105 answers, Claude 151.
ChatGPT and Grok tie on the composite at 84.4 — by very different routes: ChatGPT on links and prices, Grok on filters and deals.
Filter respect never drops below 92 and Sources never below 83. The floor of agentic shopping is already high.
Buy path runs from 27 to 75 — the difference between naming a product and getting you to a checkout.
Read across 3,112 answers, each one shops in its own recognisable way — and each is best at something.
{{ d.sig }}
3,231 of 8,471 product mentions across the 11 shopping aisles went to home-grown brands — five in six mattress picks, three-quarters of sunscreens and office chairs, none of the phones or laptops. Read each bar as: of every 100 products named in this aisle, this many were Indian.
“Named #1” counts the answers that put the brand in the top spot. Green Soul (chairs) and Minimalist (skincare) sit beside Samsung at the very top — two Indian D2C brands that the assistants now treat as default answers.
1,265 mismatched quotes where we could compare both numbers. The median miss was 8%. More than half the misses were under the shelf price — the assistant remembered a sale, not a mark-up.
If a quote {{ tolLabel }} is good enough for you, then {{ good }} of the 8,309 AI price quotes we checked would have served you well.
Exact matches (73.4%) plus the mismatches inside your tolerance, assuming the 945 mismatches without a comparable pair follow the same distribution as the 1,265 that had one.
We pinned runs to Mumbai, the National Capital Region and Bengaluru, asked every question again in Hinglish, and asked it as three different shoppers. The shortlist changed with the shopper — which is exactly what a local shopping assistant should do.
of products in Mumbai, NCR and Bengaluru answers were new versus the all-India runs — the shelf shifts even inside the country.
of Hinglish re-asks replaced every product; 77% kept at least one pick. Every assistant answered in Hinglish without breaking — language is not the barrier.
Thinking mode vs instant mode replaced every product in 22% of 1,936 pairs — the slower mode is not a different shopper, just a slower one. AI Overviews has no thinking mode at all.
Currency drift. When asked for a pick under a rupee budget, 7% of answers slipped into another currency — 16% for Grok, 10% for Meta AI, 2% for Claude. The one place geography still leaks.
Personas. Say who you are and the list changes: asked as an adult man, a wife or a senior citizen, every product was replaced in 33% of 7,780 pairs — 24%, 29% and 48% respectively. The senior-citizen persona is the single strongest lever we found.
Eight capturers in four Indian cities, nine weeks of captures (29 June – 31 August 2026), and a second layer of humans checking the first.
93 real questions across 12 categories — a budget ladder at Indian price points, then the use-cases people shop with. Asked on the assistants' current builds, geo-pinned, week after week.
Up to three merchants per product, opened by a human: 16,497 pages so far. Dead pages, homepage redirects, wrong items and out-of-stock listings all go on the record.
Prices, ratings and offers are compared to the live page, screenshot against screenshot. Then 22,894 probes: the same question in Hinglish, as a different person, in thinking vs instant mode, and again a week later.
Independent raters re-score captured answers without seeing each other. Disagreements become disputes; disputes are adjudicated under a versioned codebook; contested claims are re-run.
Two independent raters re-scored 342 captured answers field by field — 29,781 judgments. They agreed 78% of the time and said "can't tell from the evidence" on 8.5%.
Each category gets a budget ladder at real Indian price points, then the use-cases people actually shop with. Pick an aisle to read every question exactly as it was typed into the six assistants. Wave 2 adds Perplexity, deeper city pinning and a second Hinglish pass.
2,596 paired re-asks — same assistant, same words, same week. 4% came back identical; 3 in 4 kept at least one pick in common. Then we changed one thing at a time.
Tell the assistant you're a senior citizen and half the shortlist is replaced — the single strongest lever we found. Personalisation is real, and it is working.
Real picks from ChatGPT's 16 answers to this question. Outlined rows carried over from the first answer. Across the field only 4% of same-week re-asks came back identical; Grok was the steadiest, returning the same set of three 39% of the time.
of budget-capped answers stayed under the cap — 2,485 checks. 11% went over budget, and 7% quietly switched to another currency.
of cited offers were exactly right at the merchant, 17% partly right, 32% wrong — 4,340 checked. Grok (69%) and ChatGPT (67%) cite the most trustworthy deals.