Research

Which sources do AI engines cite for shopping questions? (September 2026)

Published · Data collected

On September 4, 2026 PositionBird sent 36 buyer-intent shopping questions, in 6 product categories, to 3 AI engines through their official APIs, 3 times each, and recorded every URL the answers cited. The 324 calls produced 4,626 citations to 1,328 distinct domains; this page reports which domains and which kinds of source those citations went to, and how much they change from one answer to the next.

324

API calls (318 returned an answer)

4,626

cited URLs recorded

1,328

distinct domains cited

62%

of 107 prompt-engine pairs re-cited a first-sample source in every later sample

PositionBird sells AI-visibility measurement, so it has an interest in this topic; the method, the exclusions and the raw data are published below so the figures can be checked. Engine and company names are used to identify the services measured. PositionBird is not affiliated with, endorsed by or sponsored by OpenAI, Anthropic or Google. All answers analysed were generated by AI models.

How was this measured?

Three engines were queried through their official developer APIs with web search or grounding turned on, from one account each, within the vendor's rate limits, and never through a consumer chat interface:

  • OpenAI: gpt-4.1-mini (Responses API + web_search)
  • Anthropic: claude-opus-4-8 (Messages API + web search)
  • Google Gemini: gemini-3.6-flash (Google Search grounding)

The prompt set is 36 buyer-intent questions, 6 in each of 6 product categories (electronics accessories, home goods, outdoor gear, pet supplies, skincare, supplements), in four shapes: Best-of (“best X for Y”, “top rated X”); Comparison (“A vs B”); Worth-it (“is X worth it”); Where-to-buy (“where to buy X”). Prompts were sent verbatim, in English, with no location or persona context. Each prompt-engine pair was queried 3 times, for 324 calls in total, between and . A citation is one URL in the answer's source list; the same domain cited twice in one answer counts twice.

Anthropic's run had 6 failed calls out of 108, so its figures cover 35 of the 36 prompts. A failed call is recorded as failed and excluded from every rate on this page: it is not counted as an answer with zero citations. Per-answer figures for a partially covered engine are computed over the answers it returned, and its domain counts are lower than they would be with full coverage, so cross-engine count comparisons should be read with the coverage column in mind.

Exclusions, as recorded in the dataset: Perplexity (its terms restrict publishing outputs) and Google AI Overviews (no official API) were not collected. Gemini grounding redirect URLs were resolved to their target hosts before counting.

The Gemini API returns citations as Google redirect links rather than destination URLs; each one was followed to its destination host before counting. The searched figure records whether the engine actually ran a web search for that answer. The OpenAI and Anthropic APIs were called with the search tool enabled on every request; Gemini's grounding is dynamic, so the model decides per request whether to search, and an answer without a search has no citations by construction.

API answers are a proxy for what a shopper sees in the ChatGPT, Claude or Gemini apps, not a copy of them: the consumer apps may use different model versions, system prompts, retrieval, memory and personalisation. PositionBird's standing measurement rules, which this study follows, are on the methodology page.

How many sources does each engine cite?

Per-engine citation counts: answers returned, citations per answer, share of answers with any citation, share of answers where the engine searched
EngineModelAnswersPrompts coveredCitations / answerAnswers with a citationAnswers that searched
OpenAIgpt-4.1-mini (Responses API + web_search)108 of 10836 of 366.0577.8% (84)100.0% (108)
Anthropicclaude-opus-4-8 (Messages API + web search)102 of 10835 of 3635.96100.0% (102)100.0% (102)
Google Geminigemini-3.6-flash (Google Search grounding)108 of 10836 of 362.8233.3% (36)33.3% (36)

OpenAI returned 108 answers with 653 citations, 6.05 per answer; 77.8% of its answers cited at least one source and 100.0% ran a web search.

Anthropic returned 102 answers with 3,668 citations, 35.96 per answer; 100.0% of its answers cited at least one source and 100.0% ran a web search.

Google Gemini returned 108 answers with 305 citations, 2.82 per answer; 33.3% of its answers cited at least one source and 33.3% ran a web search; the 72 answers that did not search carry no citations.

Which sources are cited most?

The 25 most-cited domains across all 4,626 citations. Share is the domain's citations divided by all citations; engines lists which engines cited it at least once; prompts is how many of the 36 prompts it appeared under.

Top 25 cited domains with category, citation count, share of all citations, engines and prompts
#DomainCategoryCitationsShareEnginesPrompts
1walmart.comMarketplace3026.5%OpenAI, Anthropic30
2ebay.comMarketplace1563.4%OpenAI, Anthropic28
3skinsort.comEditorial1042.3%Anthropic6
4amazon.comMarketplace992.1%Anthropic21
5forbes.comEditorial932.0%OpenAI, Anthropic, Google Gemini20
6youtube.comCommunity901.9%OpenAI, Anthropic, Google Gemini27
7ebay.deMarketplace731.6%Anthropic21
8healthline.comEditorial591.3%OpenAI, Anthropic, Google Gemini11
9nbcnews.comEditorial531.1%OpenAI, Anthropic, Google Gemini13
10tomsguide.comEditorial521.1%OpenAI, Anthropic, Google Gemini11
11slickdeals.netCommunity471.0%Anthropic20
12techradar.comEditorial320.7%OpenAI, Anthropic, Google Gemini4
13rei.comRetailer310.7%OpenAI, Anthropic, Google Gemini7
14cleverhiker.comEditorial300.7%OpenAI, Anthropic, Google Gemini6
15outdoorgearlab.comEditorial290.6%Anthropic, Google Gemini6
16sleepfoundation.orgEditorial280.6%OpenAI, Anthropic, Google Gemini3
17chewy.comRetailer270.6%OpenAI, Anthropic, Google Gemini5
18en.wikipedia.orgEditorial260.6%Anthropic11
19cats.comEditorial260.6%Anthropic, Google Gemini3
20cnn.comEditorial250.5%Anthropic9
21gulfnews.comEditorial250.5%Anthropic9
22zooplus.comEditorial240.5%Anthropic4
23treelinereview.comEditorial230.5%Anthropic, Google Gemini8
24bestbuy.comRetailer220.5%OpenAI, Anthropic, Google Gemini6
25gearjunkie.comEditorial220.5%Anthropic, Google Gemini6

The top 25 domains account for 32.4% of citations; the remaining 1,303 domains share the rest, and 28 of the top 60 in the dataset were cited by only one engine.

Do the engines cite different sources?

Each engine's ten most-cited domains, with the count and its share of that engine's citations.

OpenAI

653 citations

  1. forbes.com 18 · 2.8%
  2. walmart.com 16 · 2.5%
  3. healthline.com 15 · 2.3%
  4. youtube.com 14 · 2.1%
  5. homedepot.com 11 · 1.7%
  6. ebay.com 11 · 1.7%
  7. tomsguide.com 8 · 1.2%
  8. allure.com 8 · 1.2%
  9. techradar.com 7 · 1.1%
  10. goodhousekeeping.com 7 · 1.1%

Anthropic

3,668 citations

  1. walmart.com 286 · 7.8%
  2. ebay.com 145 · 4.0%
  3. skinsort.com 104 · 2.8%
  4. amazon.com 99 · 2.7%
  5. ebay.de 73 · 2.0%
  6. forbes.com 67 · 1.8%
  7. nbcnews.com 50 · 1.4%
  8. slickdeals.net 47 · 1.3%
  9. tomsguide.com 43 · 1.2%
  10. healthline.com 40 · 1.1%

Google Gemini

305 citations

  1. youtube.com 67 · 22.0%
  2. reddit.com 10 · 3.3%
  3. rei.com 8 · 2.6%
  4. forbes.com 8 · 2.6%
  5. cleverhiker.com 7 · 2.3%
  6. techradar.com 6 · 2.0%
  7. gearjunkie.com 6 · 2.0%
  8. outdoorgearlab.com 6 · 2.0%
  9. nymag.com 5 · 1.6%
  10. pcmag.com 4 · 1.3%

2 domains appear in all 3 engines' top-15 lists: forbes.com, healthline.com. Pairwise, OpenAI and Anthropic share 5 of their top 15, OpenAI and Google Gemini share 4, and Anthropic and Google Gemini share 3. On this sample the engines' most-cited sources overlap only partly, which means a store's citation picture on one engine says little about the others.

What kinds of sources win?

Every cited domain was assigned one category by PositionBird's source classifier, the same rules the app uses for its citation ledger. Counts are citations, not domains.

Citations by source category, with per-engine counts
CategoryCitationsShareOpenAIAnthropicGoogle Gemini
Editorial3,46574.9%5642,699202
Marketplace67114.5%336380
Community2104.5%1411581
Retailer1693.6%4210918
Brand store611.3%0574
Institutional501.1%0500
Editorial
Publications, review sites, buying guides and blogs. This is also the default for any domain the classifier does not recognise, so it is an upper bound.
Marketplace
Multi-seller marketplaces where a merchant can list products directly (for example the Amazon, eBay, Walmart, Etsy and Alibaba word marks).
Retailer
Retail chains that sell many brands through a wholesale relationship rather than open listing (for example Best Buy, Target, Home Depot, Chewy, IKEA).
Brand store
A manufacturer's or brand's own storefront: the pages a store owns and controls.
Community
Forums, video and social platforms, and deal communities (for example Reddit, YouTube, Slickdeals).
Institutional
Government, universities, hospitals, standards bodies and other sources that cite but cannot be pitched.
Unclassified (none recorded)
Hosts the classifier could not read at all.

Classification is automatic and list-based. A domain not on any list is counted as editorial, so the editorial share (74.9%) is an upper bound and a small number of brand storefronts or retailers are likely inside it; the brand-store share (1.3%) is correspondingly a lower bound. The raw rows carry the domains, so any other classification can be applied to them.

Does it change by product category?

The most-cited domains in each of the 6 product categories, across all engines. Citation totals differ by category partly because of the engines' coverage, so compare the lists, not the totals.

Electronics accessories

718 citations

  1. techradar.com 32 · 4.5%
  2. walmart.com 30 · 4.2%
  3. tomsguide.com 25 · 3.5%
  4. youtube.com 25 · 3.5%
  5. bestbuy.com 21 · 2.9%
  6. belkin.com 20 · 2.8%
  7. slickdeals.net 19 · 2.6%
  8. ebay.com 16 · 2.2%

Home goods

899 citations

  1. walmart.com 76 · 8.5%
  2. ebay.com 54 · 6.0%
  3. forbes.com 32 · 3.6%
  4. amazon.com 26 · 2.9%
  5. youtube.com 19 · 2.1%
  6. tomsguide.com 18 · 2.0%
  7. sleepfoundation.org 16 · 1.8%
  8. ebay.de 14 · 1.6%

Outdoor gear

696 citations

  1. rei.com 30 · 4.3%
  2. outdoorgearlab.com 29 · 4.2%
  3. ebay.com 29 · 4.2%
  4. cleverhiker.com 26 · 3.7%
  5. ebay.de 21 · 3.0%
  6. switchbacktravel.com 21 · 3.0%
  7. walmart.com 16 · 2.3%
  8. youtube.com 16 · 2.3%

Pet supplies

877 citations

  1. walmart.com 90 · 10.3%
  2. chewy.com 27 · 3.1%
  3. cats.com 26 · 3.0%
  4. amazon.com 24 · 2.7%
  5. zooplus.com 24 · 2.7%
  6. forbes.com 23 · 2.6%
  7. petco.com 21 · 2.4%
  8. ebay.com 19 · 2.2%

Skincare

775 citations

  1. skinsort.com 104 · 13.4%
  2. walmart.com 39 · 5.0%
  3. nbcnews.com 22 · 2.8%
  4. ebay.com 16 · 2.1%
  5. today.com 15 · 1.9%
  6. forbes.com 14 · 1.8%
  7. gulfnews.com 13 · 1.7%
  8. cerave.com 13 · 1.7%

Supplements

661 citations

  1. walmart.com 51 · 7.7%
  2. healthline.com 39 · 5.9%
  3. amazon.com 28 · 4.2%
  4. ebay.com 22 · 3.3%
  5. ncbi.nlm.nih.gov 19 · 2.9%
  6. forbes.com 17 · 2.6%
  7. sleepfoundation.org 12 · 1.8%
  8. health.yahoo.com 11 · 1.7%

How stable are citations from one answer to the next?

The same prompt sent to the same engine 3 times does not return the same sources. The stability metric here counts a prompt-engine pair as stable when every later sample re-cited at least one domain from the first sample; a pair needs at least two returned answers to count. Across 107 pairs, 66 were stable by that definition, 61.7%.

Citation stability by engine
EnginePairsStable pairsShare
OpenAI362672.2%
Anthropic3535100.0%
Google Gemini36513.9%

The differences between engines here track how many sources each one cites per answer: an answer with many citations has more chances to repeat one, and an answer that did not search has none. Read the engine rows alongside the citations-per-answer and searched figures above rather than as a ranking.

What this implies for measurement: a single answer is a sample, not a result. Whether a store is cited for a question is a frequency to be estimated over many samples on a schedule, with the sample counts shown, and a one-off check in any chat app, in either direction, is not evidence of much.

Does the question shape matter?

Citations per answer by prompt shape
Prompt shapeAnswersCitationsCitations / answer
Best-of (“best X for Y”, “top rated X”)1582,40715.23
Comparison (“A vs B”)5477314.31
Worth-it (“is X worth it”)5479414.70
Where-to-buy (“where to buy X”)5265212.54

Per-answer citation counts are pooled across engines, so the shape with the most answers from the heaviest-citing engine will show the highest figure; these are descriptive, and the sample per shape is small. Answers counts differ between shapes because of the prompt mix and the failed calls noted above.

What does this mean for a store?

In this sample, pages a brand owns and controls were a small share of what the engines cited (1.3% classified as brand stores), while marketplaces (14.5%) and editorial pages (74.9%, an upper bound) carried most citations. If that pattern holds for a store's own questions, being cited depends less on the store's own pages alone and more on whether the store appears on the marketplace listings, buying guides, reviews and community threads the engines already pull from. Owned pages still matter, since they are what an engine can cite directly, but presence on cited third-party pages and on marketplaces matters alongside them. Because the sources also vary from answer to answer and from engine to engine, the practical move is to measure which sources are cited for the store's own questions, repeatedly, and work on those; a citation of a page about the store without the store being named is a ghost citation, and worth tracking separately. What to monitor and how to set it up is covered in AI search monitoring for ecommerce.

Limitations

  • 36 prompts in 6 product categories, chosen by PositionBird; other prompts, categories or phrasings may cite differently.
  • One day of collection (September 4, 2026). Engines change models and retrieval without notice; these figures describe that day.
  • English-language prompts with no location or persona context, so a US-leaning result set; no other languages or markets were measured.
  • Specific models: gpt-4.1-mini (Responses API + web_search); claude-opus-4-8 (Messages API + web search); gemini-3.6-flash (Google Search grounding). Other models from the same vendors may behave differently.
  • Official APIs, not the consumer apps. What a shopper sees in a chat app can differ in model, retrieval, memory and personalisation.
  • Anthropic's run is partial: 6 of 108 calls failed, covering 35 of 36 prompts. Its per-answer rates are over the answers it returned; its domain counts are lower than a full run would give.
  • Source categories come from an automatic, list-based classifier; unknown domains default to editorial.
  • Perplexity and Google AI Overviews are not included, for the reasons stated in the exclusions above.
  • Counts are citations, not clicks or sales; a cited page is not necessarily a visited one.

Get the data

The raw rows are published as newline-delimited JSON, one object per API call (324 rows), under the Creative Commons Attribution 4.0 licence. Each row carries the timestamp, product category, prompt, engine, model identifier, sample number, whether the call succeeded, whether the engine searched, and the list of cited URLs as returned by the API. Gemini rows contain the original redirect URLs.

Download the dataset (citation-study-2026-09-04.ndjson)

Cite as: PositionBird (StatusBird LLC), Which sources do AI engines cite for shopping questions? (September 2026), September 4, 2026, https://positionbird.io/research/which-sources-ai-engines-cite-for-shopping-questions.