{{ r.label }} {{ r.n }}
01
Page 01 · Agentic commerce · why this benchmark exists

Buying through clicks is ending. Buying through words has begun.

Search, scroll, filter, compare, add to cart — for thirty years online shopping was a sequence of clicks. The next shopper simply says what they need, and an AI assistant does the rest. That shift is agentic commerce. This site is the first human-verified look at how well it already works for Indian shoppers.

Then · buying through clickssearch · scroll · filter · compare · cart
Now · buying through wordsone sentence
Best sunscreen for oily skin under ₹1,000
AI
Three picks under ₹1,000 that suit oily skin — with prices, ratings and where to buy them. No scrolling, no filters, no cart.
01

From clicks to words

Every step of online shopping was a click: search, scroll, filter, compare, cart. The next shopper says what they need — “a sunscreen for oily skin under ₹1,000” — and the assistant does the rest. The interface itself is changing.

02

Five surfaces, for the first time

For the first time in the history of commerce, buying will happen across five foundational surfaces at once — the Big 5 AI labs — and possibly more. No single storefront and no single search box owns the shopper any longer.

03

The salesman left. The paid post arrived.

E-commerce ended the manipulation of the shop-floor salesman by letting us do our own research. Then the system caught up: paid reviews, sponsored creators and generated content now shape what you think — without you knowing it.

04

An agent that is not easily swayed

We believe agentic commerce solves this. A highly intelligent entity that knows your needs, and is not moved just because your favourite content creator posted about a product or because a brand paid for a review you happened to see.

05

The beginning of the end of advertising as we know it

A language model’s natural input and output is words. For at least some product verticals, the return on visual advertising will keep falling — the shopper those ads were made for is no longer looking at them.

02
Page 02 · The Big 5 foundational labs

Five labs. Six surfaces. One shopper.

Commerce has never been contested by five foundational model makers at once. Each arrives with a different history, a different relationship with advertising — and a different reason to win the shopper.

O
OpenAI
ChatGPT
New entrant · the strongest models

The newest name in commerce and advertising — and the maker of the strongest models. It has never sold a shopping ad; it is learning the shelf in public, and fast.

G
Google
Gemini & AI Overviews
The incumbent · first real competition

The tried and trusted player: search, shopping and ads have been Google’s for two decades. For the first time in its existence it faces competition on its own ground — and it answers on two surfaces.

A
Anthropic
Claude
New entrant · consumer focus still open

A new entrant whose appetite for the consumer market is not yet clear. Yet many of us already turn to Claude for buying decisions — so it is on the shelf, whether it planned to be or not.

x
xAI
Grok
Structural entrant · compute, talent, data

Compute, talent density and data at enormous scale, plus years of advertising experience through X. Its latest models trail the Big 3 closely — and it arrived with a business model already attached.

M
Meta
Meta AI
Structural entrant · the ads machine

The largest advertising machine ever built, now shipping models that closely trail the Big 3 — with billions of shoppers already inside its apps and years of ad data behind every answer.

What we did
We are evaluating these Big 5 — on all six surfaces — against how an average consumer actually moves through a purchase.
See how we scored it
03
Page 03 · How we evaluated

We scored the journey, not the demo.

An average shopper doesn’t ask one perfect question. They ask, get filtered, weigh the options, hunt for a deal and end up on Amazon.in or Flipkart. We followed every answer through those six steps — then opened everything it pointed to.

Also analysed · not scored
Instant vs thinking mode
How far the shortlist moves when the same model is asked to think longer before answering.
Hindi prompts
How the recommendations change when the same question is asked in Hinglish — Hindi in Roman script, the way India actually types.
Personas
How the shortlist changes when the same question is asked as an adult man, as a wife shopping for the family, or as a senior citizen.
1

Understand the want

clarifying questions

Did the assistant fully understand what the shopper wanted — did it ask clarifying questions before it recommended anything, or did it guess?

2

Respect every filter

filter adherence · budget cap

Did it honour every filter in the question? “Sunscreen for oily skin within ₹1,000” carries two — skin type and price — and we checked both.

3

Explain the trade-offs

pros & cons

Did it explain the pros and cons of the shortlisted products, or simply name three and stop?

4

Find the deal — and be right about it

offers cited · offers true

Did it fetch the best offers, and when we opened the merchant page, were those offers actually there?

5

Point to Amazon.in and Flipkart

link · price match · review match

Did it hand over Amazon.in and Flipkart links — and did the prices and reviews it quoted from those two surfaces match what the pages showed?

6

Check every other shelf

page live · price truth · review truth

For every other merchant or brand link it gave, did the page open on the right product, and did the price and reviews match the brand or merchant site?

04
Page 04 · The numbers · India’s first human-verified trust benchmark for AI shopping

3,112 questions asked by hand. 88,982 receipts.

Nothing here is synthetic or self-graded. Every figure on this site traces back to a human asking a question, opening a page and writing down what they saw — in India, for Indian shoppers.

Live capture · replayed from the database {{ sceneCounter }}
{{ sceneQ }}
{{ sceneInitial }}
{{ sceneModel }} {{ sceneBuild }}
{{ pk.rank }}
{{ pk.name }}
3 links opened by hand prices & ratings vs the page asked again — in Hinglish, as someone else, a week later
Receipt · Wave 1 · India only29 Jun – 31 Aug 2026
Shopping questions asked by hand
93 questions × 6 surfaces, week after week
{{ cAsks }}
Follow-up probes
Hinglish, personas, thinking mode, budget caps, deals
{{ cProbes }}
Products named
2,858 distinct products across 11 aisles
{{ cProducts }}
Merchant pages opened by hand
every link, up to three merchants per product
{{ cPages }}
Answers double-checked blind
342 answers re-scored field by field by a second rater
{{ cRatings }}
Spot-checks
214 answers audited a third time · 98.9% upheld
{{ cChecks }}
Human-recorded data points
answers + probes + products + pages + blind ratings + spot-checks
{{ cPoints }}
93 questions · 12 categories · real rupee budgets 6 surfaces from 5 labs · exact model build logged on every answer 8 capturers · Mumbai, National Capital Region (NCR), Bengaluru, Coimbatore 4 blind raters · 5 rating rounds · codebook v0.24
Best sunscreen for oily skinBest smartphone under ₹20,000Best whey protein for sensitive stomachBest office chair for back painBest earbuds for gaming / low latencyBest cooling mattress for hot sleepersBest multivitamin for womenBest laptop for video editing
Best sunscreen for oily skinBest smartphone under ₹20,000Best whey protein for sensitive stomachBest office chair for back painBest earbuds for gaming / low latencyBest cooling mattress for hot sleepersBest multivitamin for womenBest laptop for video editing
Best fitness watch with GPS for runningBest running shoesBest sunscreen with no white cast for Indian skinBest vitamin B12 for vegetariansBest face wash for acne and pimplesBest queen mattress under ₹15,000Best mesh office chair with headrestBest plant-based protein powder
Best fitness watch with GPS for runningBest running shoesBest sunscreen with no white cast for Indian skinBest vitamin B12 for vegetariansBest face wash for acne and pimplesBest queen mattress under ₹15,000Best mesh office chair with headrestBest plant-based protein powder
05
Page 05 · What already works

Most of the time, the assistant does exactly what you asked.

Across 3,112 base answers and 22,894 follow-up probes, the basics of agentic shopping are already solid: constraints respected, budgets honoured, prices quoted, buy links handed over. Each ring is a share of all answers.

3 concrete picks in 90% of answers · 8% asked back instead 80% explained the trade-offs 95% handed over an Amazon.in or Flipkart link
{{ w.big }}
{{ w.label }}
9 in 10

links from the best assistant opened on a live product page — ChatGPT's 91%. The shelf is reachable when the model is careful.

98–100

Source scores for Gemini, Claude, AI Overviews, Grok and Meta AI: the pages they cite exist and resolve. Citation hygiene is largely solved.

58%

of price misses were within 10% of the shelf price; fewer than 3% were off by more than half. When AI gets the price wrong, it is usually a little wrong.

06
Page 06 · The shelf check

Two in three links land on a live shelf. Here's where to look twice.

16,497 merchant pages, opened by hand and compared to what the assistant claimed. Each dot is one percent of them.

67% of linked pages were live and showed the product
live page homepage redirect dead page no URL given wrong product failed, unclassified
73%

of quoted prices matched the merchant page — 6,099 of 8,309 checked. ChatGPT reached 86%; Grok 79%. When they missed, the median gap was 8%.

61%

of quoted ratings and review counts matched — 3,917 of 6,412. Claude led at 87%; AI Overviews trailed at 29%. Ratings are the claim to double-check.

49% · 60%

Amazon.in and Flipkart price quotes that matched. Marketplace prices move daily — the one number to re-check before checkout.

Claim vs shelf · drag the handle · real rows from the database
{{ clIdx }}
What {{ clModel }} said
{{ clProduct }}
at {{ clMerchant }}
What the shelf showed
{{ clShelf }}
{{ clGap }}
What {{ clModel }} said
{{ clProduct }}
at {{ clMerchant }}
Price as quoted
{{ clQuoted }}
drag right to reveal the shelf →
Why 5,232 links failed
{{ r.label }}
{{ r.n }}
Safest aisles first · pages live by category
{{ c.label }}
{{ c.v }}
07
Page 07 · Category champions

Every aisle has a champion.

Reliability is not one number — it depends on what you're buying and who you ask. Skincare and supplements are already safe ground; electronics is where the shelf moves fastest.

{{ h.name }}
{{ row.label }}
{{ c.v }}
90%+ 75–89% 60–74% 45–59% below 45% faded: fewer than 20 checks {{ matNote }}
Sunscreen · ChatGPT
98% live

The most reliable pairing in the study — 455 pages, with 94% of ratings true as well.

Whey protein · Claude
95% reviews true

Claude's supplement answers are precise: 92% of pages live, 85% of prices matched.

Office chairs · Grok
97% prices true

Grok quotes chair prices almost perfectly (36 checks); Gemini's chair ratings were 85% true.

Laptops · everyone
57% live

The hardest aisle: SKUs churn weekly. Even here ChatGPT held 88% — the gap is closable.

08
Page 08 · One question, six answers

Same words. Six shopping lists.

The products each assistant named most often for the same question — and how often it reached for them. Different assistants, genuinely different taste.

{{ c.initial }}
{{ c.model }} {{ c.runs }}
{{ p.name }}
{{ p.n }}

Counts are how many of that assistant's answers to this question named the product. Claude and Grok were captured on fewer runs in Wave 1. Auto-advances every few seconds; hover to pause.

09
Page 09 · The leaderboard

Who to lean on — and for what.

{{ dimDesc }}

Composite, Truth, Buy path, Filter and Sources come from the metrics pipeline (codebook v0.24, Wave 1 India). Everything else is computed straight from the India capture database — 3,112 answers, 16,497 merchant pages, 4,340 deal checks. † thin sample — Grok 105 answers, Claude 151.

{{ r.rank }}
{{ r.model }}
{{ r.build }}
{{ r.val }}
Joint first

ChatGPT and Grok tie on the composite at 84.4 — by very different routes: ChatGPT on links and prices, Grok on filters and deals.

Everyone passes

Filter respect never drops below 92 and Sources never below 83. The floor of agentic shopping is already high.

Biggest spread

Buy path runs from 27 to 75 — the difference between naming a product and getting you to a checkout.

10
Page 10 · Six personalities · keep scrolling, the cards stack

Every assistant has a strength worth knowing.

Read across 3,112 answers, each one shops in its own recognisable way — and each is best at something.

{{ d.initial }}
{{ d.model }}
{{ d.build }} · {{ d.runs }} answers · dossier {{ d.idx }}
{{ d.best }}

{{ d.sig }}

{{ s.k }}
{{ s.v }}
11
Page 11 · The marketplace the assistants built · scroll down to travel sideways

Twenty shelves carry the traffic. The specialists deliver.

Every merchant assistants sent shoppers to at least 175 times, with how often its page opened and how often the quoted price held. Nykaa, 1mg, Purplle and Myntra — the category specialists — are the most dependable landing spots.

{{ m.idx }}{{ m.links }}
{{ m.name }}
{{ m.live }}
{{ m.liveLabel }}
{{ m.price }}
prices matched

{{ m.note }}

12
Page 12 · The brands AI reaches for

More than 1 in 3 products AI recommends to Indian shoppers is an Indian brand.

3,231 of 8,471 product mentions across the 11 shopping aisles went to home-grown brands — five in six mattress picks, three-quarters of sunscreens and office chairs, none of the phones or laptops. Read each bar as: of every 100 products named in this aisle, this many were Indian.

Aisle by aisle · share of mentions that were Indian brands · and the three brands named most
Indian brandglobal brand
{{ b.label }}
{{ b.inLabel }}
{{ b.v }}
{{ t.name }} {{ t.pct }}
The ten brands AI names most · across 8,480 product mentions
#BrandMentionsNamed #1Home aisleOrigin
{{ b.rank }} {{ b.name }}
{{ b.n }} {{ b.r1 }} answers {{ b.aisle }} {{ b.tag }}

“Named #1” counts the answers that put the brand in the top spot. Green Soul (chairs) and Minimalist (skincare) sit beside Samsung at the very top — two Indian D2C brands that the assistants now treat as default answers.

13
Page 13 · How wrong is wrong?

When AI misquotes a price, it is usually off by a coffee, not a phone.

1,265 mismatched quotes where we could compare both numbers. The median miss was 8%. More than half the misses were under the shelf price — the assistant remembered a sale, not a mark-up.

{{ b.label }}
{{ b.n }}{{ b.pct }}
53%
quoted below the shelf price
47%
quoted above it
8%
median gap when wrong
Set your own tolerance · slide
{{ good }}good enough

If a quote {{ tolLabel }} is good enough for you, then {{ good }} of the 8,309 AI price quotes we checked would have served you well.

exact only · 73%within 10% · 89%within 50% · 99%

Exact matches (73.4%) plus the mismatches inside your tolerance, assuming the 945 mismatches without a comparable pair follow the same distribution as the 1,265 that had one.

14
Page 14 · Language and city

Same question, a different city or a different language — the shelf moves with the shopper.

We pinned runs to Mumbai, the National Capital Region and Bengaluru, asked every question again in Hinglish, and asked it as three different shoppers. The shortlist changed with the shopper — which is exactly what a local shopping assistant should do.

Pinned to a city · 210 answers
48%

of products in Mumbai, NCR and Bengaluru answers were new versus the all-India runs — the shelf shifts even inside the country.

Asked in Hinglish · 2,588 pairs
23%

of Hinglish re-asks replaced every product; 77% kept at least one pick. Every assistant answered in Hinglish without breaking — language is not the barrier.

22%

Thinking mode vs instant mode replaced every product in 22% of 1,936 pairs — the slower mode is not a different shopper, just a slower one. AI Overviews has no thinking mode at all.

7%

Currency drift. When asked for a pick under a rupee budget, 7% of answers slipped into another currency — 16% for Grok, 10% for Meta AI, 2% for Claude. The one place geography still leaks.

33%

Personas. Say who you are and the list changes: asked as an adult man, a wife or a senior citizen, every product was replaced in 33% of 7,780 pairs — 24%, 29% and 48% respectively. The senior-citizen persona is the single strongest lever we found.

15
Page 15 · How we know

No synthetic tasks. No self-grading. Every number above has a receipt.

Eight capturers in four Indian cities, nine weeks of captures (29 June – 31 August 2026), and a second layer of humans checking the first.

1

Ask like a shopper

93 real questions across 12 categories — a budget ladder at Indian price points, then the use-cases people shop with. Asked on the assistants' current builds, geo-pinned, week after week.

2

Open every link

Up to three merchants per product, opened by a human: 16,497 pages so far. Dead pages, homepage redirects, wrong items and out-of-stock listings all go on the record.

3

Check the shelf, then push

Prices, ratings and offers are compared to the live page, screenshot against screenshot. Then 22,894 probes: the same question in Hinglish, as a different person, in thinking vs instant mode, and again a week later.

4

Weigh it blind

Independent raters re-score captured answers without seeing each other. Disagreements become disputes; disputes are adjudicated under a versioned codebook; contested claims are re-run.

78%rater agreement

Two independent raters re-scored 342 captured answers field by field — 29,781 judgments. They agreed 78% of the time and said "can't tell from the evidence" on 8.5%.

8,218
field-level spot-checks on 214 answers · 98.9% upheld
9.3 s
average deliberation per blind judgment · 72 rater-hours in total
57
disputes raised and adjudicated · 11 claims escalated for re-run
8 capturers · Mumbai, National Capital Region (NCR), Bengaluru, Coimbatore 4 blind raters · 5 rating rounds codebook v0.24 · every edit audit-trailed exact model build logged on every answer 88,982 = 3,112 answers + 22,894 probes + 8,480 products + 16,497 pages + 29,781 blind ratings + 8,218 spot-checks geo-pinned: India, plus Mumbai, NCR and Bengaluru runs
16
Page 16 · Coverage · every question we asked

93 questions. 12 categories. Real rupees.

Each category gets a budget ladder at real Indian price points, then the use-cases people actually shop with. Pick an aisle to read every question exactly as it was typed into the six assistants. Wave 2 adds Perplexity, deeper city pinning and a second Hinglish pass.

Now showing
{{ qCat }}
{{ qCount }}
{{ q.id }}
{{ q.text }}
{{ q.intent }}
17
Page 17 · The frontier

Ask the same question twice and the shortlist moves. Consistency is the next thing to win.

2,596 paired re-asks — same assistant, same words, same week. 4% came back identical; 3 in 4 kept at least one pick in common. Then we changed one thing at a time.

{{ g.big }}
{{ g.label }}
{{ g.sub }}
Say who you are and the list changes · share of pairs where every product changed
{{ p.label }}
{{ p.v }}

Tell the assistant you're a senior citizen and half the shortlist is replaced — the single strongest lever we found. Personalisation is real, and it is working.

"Best earbuds under ₹3,000" ChatGPT · 16 answers
First answer
{{ b.rank }}{{ b.name }}
{{ altLabel }}
{{ a.rank }}{{ a.name }}
{{ a.rank }}{{ a.name }}
{{ a.rank }}{{ a.name }}

Real picks from ChatGPT's 16 answers to this question. Outlined rows carried over from the first answer. Across the field only 4% of same-week re-asks came back identical; Grok was the steadiest, returning the same set of three 39% of the time.

"Under ₹X?"
81%

of budget-capped answers stayed under the cap — 2,485 checks. 11% went over budget, and 7% quietly switched to another currency.

"Any deals on that?"
49%

of cited offers were exactly right at the merchant, 17% partly right, 32% wrong — 4,340 checked. Grok (69%) and ChatGPT (67%) cite the most trustworthy deals.

The Wave 1 report · for brand and merchant teams

To get into more details about the report, please enroll.

Enroll with your official brand or merchant email. We will walk you through how the six assistants treat your category, which merchants they trust, where your brand already appears — and what the full 88,982-point dataset can tell your team.

Category deep-dives Merchant trust map Brand-level cuts · weekly tracking
I work at a
Official work email
{{ fieldMsg }}

Brand, merchant and agency enrollments need an official work email. If you are simply interested in knowing more, enroll under Other — any email works there. A human replies, not a drip campaign.

You’re enrolled.

We’ll reach out to {{ joinedEmail }} with the {{ joinedRole }} walkthrough of the Wave 1 report — your category, the merchants the assistants trust, and where your brand already appears.