Why ChatGPT and Gemini Give Different Shopping Advice

A September 2026 study found ChatGPT and Gemini rarely cite the same shopping sources, raising major questions about AI visibility tracking.

A September 2026 preprint found that ChatGPT and Gemini often cite completely different sources for the same shopping questions, with just 5.4% average source overlap. The study also found significant variation between repeated queries and consumer apps versus APIs. For marketers, this means a single AI visibility check can be misleading, making repeated, controlled measurements essential for evaluating whether a brand is consistently recommended.

Here’s something that should worry anyone tracking “AI visibility” for a brand: according to a new September 2026 preprint, when ChatGPT and Gemini answered the same product question, they showed zero overlapping source websites in over three-quarters of matched comparisons. Not a different ranking of the same sites — completely different sites.

I’ve spent the last few years building reporting dashboards for clients who want to know “are we showing up in ChatGPT?” This study is basically a documented version of the thing I’ve been telling them anecdotally: a single screenshot of an AI answer tells you almost nothing repeatable.

Why it matters

If you’re running content strategy, SEO, or PR and you’ve started tracking “AI search visibility” as a KPI, this changes what that metric can honestly claim to measure. A brand mention in one ChatGPT session isn’t evidence of consistent placement — it might not even show up the same way five minutes later, on the same platform, from the same question.

For marketers, that means the difference between “we got cited once” and “we’re reliably recommended” is enormous, and most visibility tools blur that line right now. For publishers and retailers, it means chasing a single favorable AI answer is a much shakier strategy than it looks in a slide deck.

What the researchers actually tested

The team built a dataset called ConsumerQ from over 2,500 real commercial-advice questions, then hand-picked 117 broad product questions for controlled testing. Each question ran three times across five setups: the logged-out ChatGPT and Gemini consumer apps, their respective APIs, and Google AI Overviews. That produced roughly 1,536 usable answers.

The testing ran September 4–5, 2026, from residential Dutch IP addresses with fresh browser profiles in English. Worth flagging up front — this describes that specific test environment, not a global average across every account type, region, or future model update. Anyone quoting this as “AI search behaves like X” everywhere is overstating it.

The numbers: source overlap is the real story

The headline stat is stark: average source-domain overlap between ChatGPT and Gemini answers to the same question was just 5.4%, and 76.7% of matched checks shared zero displayed source domains.

It gets messier when you look at consistency within a single platform. Repeating the same question three times, average domain overlap was 26.0% for ChatGPT, 29.8% for Gemini, and 45.9% for Google AI Overviews. Product overlap across repeats followed a similar pattern — AI Overviews was the most stable of the three, ChatGPT the least.

I’ve seen this firsthand running the same prompt against ChatGPT twice in a row for a client audit and getting genuinely different brand names back. It’s easy to write that off as a fluke. This study suggests it’s closer to the norm.

The API-versus-app problem nobody’s dashboard accounts for

This is the part I’d flag hardest to anyone buying an AI-tracking tool. The researchers also compared the consumer apps against their APIs, and found average source overlap of only 12.0% for ChatGPT and 14.8% for Gemini — with no shared domain at all in 60.9% of ChatGPT pairs and 43.4% of Gemini pairs.

Here’s the catch, and the study is upfront about it: this doesn’t prove the interface alone is causing the gap. The exact logged-out ChatGPT model version couldn’t be confirmed, the API call was an approximation of it, and retrieval settings differ too. So it’s not a clean cause, but it’s still a real warning sign — a lot of visibility-tracking products run their checks through the API because it’s cheaper and more scriptable, then report those numbers as if they represent what a shopper actually saw in the app. Based on this data, that’s a meaningfully different thing.

“I would buy this” isn’t the same as testing it

One stylistic detail stood out to me: ChatGPT used confident, first-person recommendation language (“I would buy…”) in 79% of product answers, versus just 7% for Gemini and 2% for AI Overviews. That tone difference says nothing about which system actually evaluated the products more rigorously — it’s a writing style choice, not evidence of hands-on testing.

For anyone reading these answers as a consumer, or building content meant to influence them, that distinction matters. A confident-sounding recommendation and an actual product comparison are not the same artifact.

Limitations worth sitting with

This is a preprint, and the 117 product questions were selected for the analysis rather than sampled to represent all shoppers. It measures the tested engines, in that country, in that window, with that configuration — not a permanent statement about how these assistants behave everywhere, forever. Model updates alone could shift these numbers within weeks.

It also measures displayed sources and named products, not clicks, conversions, or actual purchases. Showing up in an answer and driving a sale are two different claims, and the study doesn’t try to connect them.

What this means if you’re actually trying to track this stuff

The researchers’ suggested audit method is basically what I’d tell a client to do anyway, and it’s more disciplined than most agencies bother with: pick around ten genuinely different buying questions (not ten phrasings of one question), run each three times per platform over several days, keep everything else fixed — country, language, device, account state — and log both the displayed domain and where the link actually resolves.

Then separate four things that keep getting collapsed into one number: a brand being mentioned, a product being recommended, a source being cited, and a visit actually being referred. Those are not interchangeable, and treating them as one “AI visibility score” is how you end up reporting a win that isn’t there.

AI shopping visibility dashboard showing inconsistent sources and recommendations

The bigger picture

None of this means AI visibility tracking is worthless — it means it needs the same rigor we eventually forced onto SEO reporting after years of vanity metrics. Right now, most tools in this space are still selling the equivalent of a single rank-tracker screenshot. This study is a decent nudge toward something closer to a proper, repeatable measurement discipline: multiple runs, labeled collection surfaces, and honesty about what a single captured answer can and can’t tell you.

If you’re evaluating an AI-search-visibility tool for your own stack, the one question worth asking the vendor is simple: how many times do they repeat a query before reporting a result, and do they tell you when the app and the API disagree? Based on what this preprint found, if the answer is “once, through the API,” take the numbers with a serious grain of salt.