324
API calls (318 returned an answer)
4,626
cited URLs recorded
1,328
distinct domains cited
62%
of 107 prompt-engine pairs re-cited a first-sample source in every later sample
PositionBird sells AI-visibility measurement, so it has an interest in this topic; the method, the exclusions and the raw data are published below so the figures can be checked. Engine and company names are used to identify the services measured. PositionBird is not affiliated with, endorsed by or sponsored by OpenAI, Anthropic or Google. All answers analysed were generated by AI models.
How was this measured?
Three engines were queried through their official developer APIs with web search or grounding turned on, from one account each, within the vendor's rate limits, and never through a consumer chat interface:
- OpenAI:
gpt-4.1-mini (Responses API + web_search) - Anthropic:
claude-opus-4-8 (Messages API + web search) - Google Gemini:
gemini-3.6-flash (Google Search grounding)
The prompt set is 36 buyer-intent questions, 6 in each of 6 product categories (electronics accessories, home goods, outdoor gear, pet supplies, skincare, supplements), in four shapes: Best-of (“best X for Y”, “top rated X”); Comparison (“A vs B”); Worth-it (“is X worth it”); Where-to-buy (“where to buy X”). Prompts were sent verbatim, in English, with no location or persona context. Each prompt-engine pair was queried 3 times, for 324 calls in total, between and . A citation is one URL in the answer's source list; the same domain cited twice in one answer counts twice.
Anthropic's run had 6 failed calls out of 108, so its figures cover 35 of the 36 prompts. A failed call is recorded as failed and excluded from every rate on this page: it is not counted as an answer with zero citations. Per-answer figures for a partially covered engine are computed over the answers it returned, and its domain counts are lower than they would be with full coverage, so cross-engine count comparisons should be read with the coverage column in mind.
Exclusions, as recorded in the dataset: Perplexity (its terms restrict publishing outputs) and Google AI Overviews (no official API) were not collected. Gemini grounding redirect URLs were resolved to their target hosts before counting.
The Gemini API returns citations as Google redirect links rather than destination URLs; each one was followed to its destination host before counting. The searched figure records whether the engine actually ran a web search for that answer. The OpenAI and Anthropic APIs were called with the search tool enabled on every request; Gemini's grounding is dynamic, so the model decides per request whether to search, and an answer without a search has no citations by construction.
API answers are a proxy for what a shopper sees in the ChatGPT, Claude or Gemini apps, not a copy of them: the consumer apps may use different model versions, system prompts, retrieval, memory and personalisation. PositionBird's standing measurement rules, which this study follows, are on the methodology page.
How many sources does each engine cite?
| Engine | Model | Answers | Prompts covered | Citations / answer | Answers with a citation | Answers that searched |
|---|---|---|---|---|---|---|
| OpenAI | gpt-4.1-mini (Responses API + web_search) | 108 of 108 | 36 of 36 | 6.05 | 77.8% (84) | 100.0% (108) |
| Anthropic | claude-opus-4-8 (Messages API + web search) | 102 of 108 | 35 of 36 | 35.96 | 100.0% (102) | 100.0% (102) |
| Google Gemini | gemini-3.6-flash (Google Search grounding) | 108 of 108 | 36 of 36 | 2.82 | 33.3% (36) | 33.3% (36) |
OpenAI returned 108 answers with 653 citations, 6.05 per answer; 77.8% of its answers cited at least one source and 100.0% ran a web search.
Anthropic returned 102 answers with 3,668 citations, 35.96 per answer; 100.0% of its answers cited at least one source and 100.0% ran a web search.
Google Gemini returned 108 answers with 305 citations, 2.82 per answer; 33.3% of its answers cited at least one source and 33.3% ran a web search; the 72 answers that did not search carry no citations.
Which sources are cited most?
The 25 most-cited domains across all 4,626 citations. Share is the domain's citations divided by all citations; engines lists which engines cited it at least once; prompts is how many of the 36 prompts it appeared under.
| # | Domain | Category | Citations | Share | Engines | Prompts |
|---|---|---|---|---|---|---|
| 1 | walmart.com | Marketplace | 302 | 6.5% | OpenAI, Anthropic | 30 |
| 2 | ebay.com | Marketplace | 156 | 3.4% | OpenAI, Anthropic | 28 |
| 3 | skinsort.com | Editorial | 104 | 2.3% | Anthropic | 6 |
| 4 | amazon.com | Marketplace | 99 | 2.1% | Anthropic | 21 |
| 5 | forbes.com | Editorial | 93 | 2.0% | OpenAI, Anthropic, Google Gemini | 20 |
| 6 | youtube.com | Community | 90 | 1.9% | OpenAI, Anthropic, Google Gemini | 27 |
| 7 | ebay.de | Marketplace | 73 | 1.6% | Anthropic | 21 |
| 8 | healthline.com | Editorial | 59 | 1.3% | OpenAI, Anthropic, Google Gemini | 11 |
| 9 | nbcnews.com | Editorial | 53 | 1.1% | OpenAI, Anthropic, Google Gemini | 13 |
| 10 | tomsguide.com | Editorial | 52 | 1.1% | OpenAI, Anthropic, Google Gemini | 11 |
| 11 | slickdeals.net | Community | 47 | 1.0% | Anthropic | 20 |
| 12 | techradar.com | Editorial | 32 | 0.7% | OpenAI, Anthropic, Google Gemini | 4 |
| 13 | rei.com | Retailer | 31 | 0.7% | OpenAI, Anthropic, Google Gemini | 7 |
| 14 | cleverhiker.com | Editorial | 30 | 0.7% | OpenAI, Anthropic, Google Gemini | 6 |
| 15 | outdoorgearlab.com | Editorial | 29 | 0.6% | Anthropic, Google Gemini | 6 |
| 16 | sleepfoundation.org | Editorial | 28 | 0.6% | OpenAI, Anthropic, Google Gemini | 3 |
| 17 | chewy.com | Retailer | 27 | 0.6% | OpenAI, Anthropic, Google Gemini | 5 |
| 18 | en.wikipedia.org | Editorial | 26 | 0.6% | Anthropic | 11 |
| 19 | cats.com | Editorial | 26 | 0.6% | Anthropic, Google Gemini | 3 |
| 20 | cnn.com | Editorial | 25 | 0.5% | Anthropic | 9 |
| 21 | gulfnews.com | Editorial | 25 | 0.5% | Anthropic | 9 |
| 22 | zooplus.com | Editorial | 24 | 0.5% | Anthropic | 4 |
| 23 | treelinereview.com | Editorial | 23 | 0.5% | Anthropic, Google Gemini | 8 |
| 24 | bestbuy.com | Retailer | 22 | 0.5% | OpenAI, Anthropic, Google Gemini | 6 |
| 25 | gearjunkie.com | Editorial | 22 | 0.5% | Anthropic, Google Gemini | 6 |
The top 25 domains account for 32.4% of citations; the remaining 1,303 domains share the rest, and 28 of the top 60 in the dataset were cited by only one engine.
Do the engines cite different sources?
Each engine's ten most-cited domains, with the count and its share of that engine's citations.
OpenAI
653 citations
- forbes.com 18 · 2.8%
- walmart.com 16 · 2.5%
- healthline.com 15 · 2.3%
- youtube.com 14 · 2.1%
- homedepot.com 11 · 1.7%
- ebay.com 11 · 1.7%
- tomsguide.com 8 · 1.2%
- allure.com 8 · 1.2%
- techradar.com 7 · 1.1%
- goodhousekeeping.com 7 · 1.1%
Anthropic
3,668 citations
- walmart.com 286 · 7.8%
- ebay.com 145 · 4.0%
- skinsort.com 104 · 2.8%
- amazon.com 99 · 2.7%
- ebay.de 73 · 2.0%
- forbes.com 67 · 1.8%
- nbcnews.com 50 · 1.4%
- slickdeals.net 47 · 1.3%
- tomsguide.com 43 · 1.2%
- healthline.com 40 · 1.1%
Google Gemini
305 citations
- youtube.com 67 · 22.0%
- reddit.com 10 · 3.3%
- rei.com 8 · 2.6%
- forbes.com 8 · 2.6%
- cleverhiker.com 7 · 2.3%
- techradar.com 6 · 2.0%
- gearjunkie.com 6 · 2.0%
- outdoorgearlab.com 6 · 2.0%
- nymag.com 5 · 1.6%
- pcmag.com 4 · 1.3%
2 domains appear in all 3 engines' top-15 lists: forbes.com, healthline.com. Pairwise, OpenAI and Anthropic share 5 of their top 15, OpenAI and Google Gemini share 4, and Anthropic and Google Gemini share 3. On this sample the engines' most-cited sources overlap only partly, which means a store's citation picture on one engine says little about the others.
What kinds of sources win?
Every cited domain was assigned one category by PositionBird's source classifier, the same rules the app uses for its citation ledger. Counts are citations, not domains.
| Category | Citations | Share | OpenAI | Anthropic | Google Gemini |
|---|---|---|---|---|---|
| Editorial | 3,465 | 74.9% | 564 | 2,699 | 202 |
| Marketplace | 671 | 14.5% | 33 | 638 | 0 |
| Community | 210 | 4.5% | 14 | 115 | 81 |
| Retailer | 169 | 3.6% | 42 | 109 | 18 |
| Brand store | 61 | 1.3% | 0 | 57 | 4 |
| Institutional | 50 | 1.1% | 0 | 50 | 0 |
- Editorial
- Publications, review sites, buying guides and blogs. This is also the default for any domain the classifier does not recognise, so it is an upper bound.
- Marketplace
- Multi-seller marketplaces where a merchant can list products directly (for example the Amazon, eBay, Walmart, Etsy and Alibaba word marks).
- Retailer
- Retail chains that sell many brands through a wholesale relationship rather than open listing (for example Best Buy, Target, Home Depot, Chewy, IKEA).
- Brand store
- A manufacturer's or brand's own storefront: the pages a store owns and controls.
- Community
- Forums, video and social platforms, and deal communities (for example Reddit, YouTube, Slickdeals).
- Institutional
- Government, universities, hospitals, standards bodies and other sources that cite but cannot be pitched.
- Unclassified (none recorded)
- Hosts the classifier could not read at all.
Classification is automatic and list-based. A domain not on any list is counted as editorial, so the editorial share (74.9%) is an upper bound and a small number of brand storefronts or retailers are likely inside it; the brand-store share (1.3%) is correspondingly a lower bound. The raw rows carry the domains, so any other classification can be applied to them.
Does it change by product category?
The most-cited domains in each of the 6 product categories, across all engines. Citation totals differ by category partly because of the engines' coverage, so compare the lists, not the totals.
Electronics accessories
718 citations
- techradar.com 32 · 4.5%
- walmart.com 30 · 4.2%
- tomsguide.com 25 · 3.5%
- youtube.com 25 · 3.5%
- bestbuy.com 21 · 2.9%
- belkin.com 20 · 2.8%
- slickdeals.net 19 · 2.6%
- ebay.com 16 · 2.2%
Home goods
899 citations
- walmart.com 76 · 8.5%
- ebay.com 54 · 6.0%
- forbes.com 32 · 3.6%
- amazon.com 26 · 2.9%
- youtube.com 19 · 2.1%
- tomsguide.com 18 · 2.0%
- sleepfoundation.org 16 · 1.8%
- ebay.de 14 · 1.6%
Outdoor gear
696 citations
- rei.com 30 · 4.3%
- outdoorgearlab.com 29 · 4.2%
- ebay.com 29 · 4.2%
- cleverhiker.com 26 · 3.7%
- ebay.de 21 · 3.0%
- switchbacktravel.com 21 · 3.0%
- walmart.com 16 · 2.3%
- youtube.com 16 · 2.3%
Pet supplies
877 citations
- walmart.com 90 · 10.3%
- chewy.com 27 · 3.1%
- cats.com 26 · 3.0%
- amazon.com 24 · 2.7%
- zooplus.com 24 · 2.7%
- forbes.com 23 · 2.6%
- petco.com 21 · 2.4%
- ebay.com 19 · 2.2%
Skincare
775 citations
- skinsort.com 104 · 13.4%
- walmart.com 39 · 5.0%
- nbcnews.com 22 · 2.8%
- ebay.com 16 · 2.1%
- today.com 15 · 1.9%
- forbes.com 14 · 1.8%
- gulfnews.com 13 · 1.7%
- cerave.com 13 · 1.7%
Supplements
661 citations
- walmart.com 51 · 7.7%
- healthline.com 39 · 5.9%
- amazon.com 28 · 4.2%
- ebay.com 22 · 3.3%
- ncbi.nlm.nih.gov 19 · 2.9%
- forbes.com 17 · 2.6%
- sleepfoundation.org 12 · 1.8%
- health.yahoo.com 11 · 1.7%
How stable are citations from one answer to the next?
The same prompt sent to the same engine 3 times does not return the same sources. The stability metric here counts a prompt-engine pair as stable when every later sample re-cited at least one domain from the first sample; a pair needs at least two returned answers to count. Across 107 pairs, 66 were stable by that definition, 61.7%.
| Engine | Pairs | Stable pairs | Share |
|---|---|---|---|
| OpenAI | 36 | 26 | 72.2% |
| Anthropic | 35 | 35 | 100.0% |
| Google Gemini | 36 | 5 | 13.9% |
The differences between engines here track how many sources each one cites per answer: an answer with many citations has more chances to repeat one, and an answer that did not search has none. Read the engine rows alongside the citations-per-answer and searched figures above rather than as a ranking.
What this implies for measurement: a single answer is a sample, not a result. Whether a store is cited for a question is a frequency to be estimated over many samples on a schedule, with the sample counts shown, and a one-off check in any chat app, in either direction, is not evidence of much.
Does the question shape matter?
| Prompt shape | Answers | Citations | Citations / answer |
|---|---|---|---|
| Best-of (“best X for Y”, “top rated X”) | 158 | 2,407 | 15.23 |
| Comparison (“A vs B”) | 54 | 773 | 14.31 |
| Worth-it (“is X worth it”) | 54 | 794 | 14.70 |
| Where-to-buy (“where to buy X”) | 52 | 652 | 12.54 |
Per-answer citation counts are pooled across engines, so the shape with the most answers from the heaviest-citing engine will show the highest figure; these are descriptive, and the sample per shape is small. Answers counts differ between shapes because of the prompt mix and the failed calls noted above.
What does this mean for a store?
In this sample, pages a brand owns and controls were a small share of what the engines cited (1.3% classified as brand stores), while marketplaces (14.5%) and editorial pages (74.9%, an upper bound) carried most citations. If that pattern holds for a store's own questions, being cited depends less on the store's own pages alone and more on whether the store appears on the marketplace listings, buying guides, reviews and community threads the engines already pull from. Owned pages still matter, since they are what an engine can cite directly, but presence on cited third-party pages and on marketplaces matters alongside them. Because the sources also vary from answer to answer and from engine to engine, the practical move is to measure which sources are cited for the store's own questions, repeatedly, and work on those; a citation of a page about the store without the store being named is a ghost citation, and worth tracking separately. What to monitor and how to set it up is covered in AI search monitoring for ecommerce.
Limitations
- 36 prompts in 6 product categories, chosen by PositionBird; other prompts, categories or phrasings may cite differently.
- One day of collection (September 4, 2026). Engines change models and retrieval without notice; these figures describe that day.
- English-language prompts with no location or persona context, so a US-leaning result set; no other languages or markets were measured.
- Specific models: gpt-4.1-mini (Responses API + web_search); claude-opus-4-8 (Messages API + web search); gemini-3.6-flash (Google Search grounding). Other models from the same vendors may behave differently.
- Official APIs, not the consumer apps. What a shopper sees in a chat app can differ in model, retrieval, memory and personalisation.
- Anthropic's run is partial: 6 of 108 calls failed, covering 35 of 36 prompts. Its per-answer rates are over the answers it returned; its domain counts are lower than a full run would give.
- Source categories come from an automatic, list-based classifier; unknown domains default to editorial.
- Perplexity and Google AI Overviews are not included, for the reasons stated in the exclusions above.
- Counts are citations, not clicks or sales; a cited page is not necessarily a visited one.
Get the data
The raw rows are published as newline-delimited JSON, one object per API call (324 rows), under the Creative Commons Attribution 4.0 licence. Each row carries the timestamp, product category, prompt, engine, model identifier, sample number, whether the call succeeded, whether the engine searched, and the list of cited URLs as returned by the API. Gemini rows contain the original redirect URLs.
Download the dataset (citation-study-2026-09-04.ndjson)
Cite as: PositionBird (StatusBird LLC), Which sources do AI engines cite for shopping questions? (September 2026), September 4, 2026, https://positionbird.io/research/which-sources-ai-engines-cite-for-shopping-questions.