Research
Which websites do AI engines cite about Shopify?
We sent the same fixed panel of merchant questions to two AI answer engines every week for seven weeks and recorded every website each one cited. The two engines agree on the questions and almost nothing else: across 405 answers they named 427 distinct domains between them, and only 12 of those appear on both lists.
- Answers measured
- 405
- Weekly waves
- 7
- Engines
- 2
- Domains seen
- 427
Window 2026-07-22 to 2026-08-28. Data last rebuilt 2026-08-29. Raw data and charts are free to reuse under CC BY 4.0.
The finding: two engines, two different internets
Both engines answered the same questions, phrased the way a merchant actually asks them. What they read to answer has almost nothing in common.
The OpenAI leg (gpt-5.4-mini with its built-in web search) drew on 32 distinct domains across 297 answers. The median answer cited 1. Shopify-owned domains appeared in 99.3% of answers, and in 80.1% of answers they were the only thing cited at all. Four answers in five contained no independent publisher of any kind.
The Gemini leg (gemini-2.5-flash with Google search grounding) drew on 407 distinct domains across 108 answers, a median of 10 per answer and up to 32. Shopify-owned domains appeared in 63.9% of answers, and in 0% of answers were the only source. Not once did this engine answer a Shopify question out of Shopify alone.
These are not two samples of one surface. Of the 427 domains the panel saw in total, exactly 12 were cited by both engines: gelato.com, printful.com, printify.com, qualimero.com, rovela.ai, shopexperts.com, shopify.com, shopify.dev, smile.io, upwork.com, woocommerce.com, yotpo.com. Strip out the platform vendors and the overlap is a handful of app and print-on-demand brands.
The practical consequence is uncomfortable for anyone tracking a single "AI visibility" score. On one of these surfaces a Shopify publisher is competing for a slot that mostly does not exist; on the other there are ten slots per answer and a field of four hundred competitors. Averaged together, those two facts produce a number that describes neither.
How this was measured
The instrument is a fixed panel of buyer-intent prompts, grouped into 5 clusters: money-trial, vs-comparison, pricing-fees, dev-services, how-to. The wording is grounded in real Search Console queries for this site rather than in keyword tools, so the prompts read like questions ("Is the Shopify one-dollar trial still active right now?") instead of search strings.
Each run sends the exact same prompt text to an engine with web access switched on and reads the citations out of the response's structured metadata: annotation objects for OpenAI, grounding chunks for Gemini. Nothing is scraped out of the prose. That distinction matters more than it sounds: asked in plain language which sources it used, an assistant will happily produce a plausible-looking list it never actually read.
The panel itself never changes between runs. That is the whole point of it, and it is also the main cost: a question we would phrase better today stays phrased the way it was in July, because the comparison across weeks is worth more than any individual improvement.
Three rules govern how the rows become numbers, and each of them changes the answer:
- Failed calls are not zeros. 84 calls in this window errored out, almost all of them the Gemini leg hitting its daily grounding quota. An engine that never answered did not decline to cite anyone. Those rows are excluded from every denominator rather than counted as a miss, which is why the two engines have very different sample sizes.
- Domains are folded to their registrable level. OpenAI reports full hostnames, so
help.shopify.comandshopify.comarrive separately. Gemini exposes only the second-level domain, so both arrive as one. Compared raw, Gemini would appear never to cite Shopify's help centre and the overlap between the engines would look smaller than it is. Every cross-engine figure here is computed on folded domains; the published CSV keeps the raw host beside the folded one so anyone can redo it the other way. - A share is only ever read inside one engine. Nothing on this page blends the two legs into a single rate. A combined figure would move whenever one engine's quota changed, which is a property of our API access, not of the web.
What these numbers cannot tell you
This section is longer than it needs to be for a marketing page and exactly as long as it needs to be for a measurement.
- The panel is not a random sample of Shopify questions. Every prompt was written to map onto a topic this site covers. It describes the slice of merchant intent we care about, and a panel built by a payments vendor or a theme shop would produce a different domain list from the same engines.
- We are in our own results. This site appears in the gemini leaderboard below. We built the instrument, we chose the prompts, and we are one of the measured parties. The raw file is published so that this is checkable rather than something you have to take on trust.
- Sample sizes are uneven. The gemini leg runs one cluster per week against a free-tier quota of roughly twenty grounded prompts a day, so its clusters were measured on different dates and one of them (how-to, n=1) barely registered in this window. Any cell below 5 answers is printed for completeness and should not be read as a rate.
- A citation is not a visit, and not a recommendation. This measures what an engine listed as a source. Whether a reader clicked it, and whether the answer above it was any good, are different questions this instrument does not touch.
- Neither leg is the Google results page. Both are chat assistants. AI Overviews and AI Mode, the surfaces most merchants actually meet, are a separate measurement we have not been able to run, and nothing here should be read as describing them.
- Engines move underneath you. Model versions are pinned in the data (gpt-5.4-mini, gemini-2.5-flash) because a provider swapping a model changes the citation behaviour without changing anything on our side. Compare like-for-like model rows, or the trend is measuring the vendor's release schedule.
How much room a publisher has
The most useful single number here is not our share. It is how often an answer contains anybody at all besides Shopify, because that is the ceiling on what any independent site can win.
Broken out by cluster, the ceiling is not uniform, and the split is by question type rather than by topic. Questions with one dated, checkable answer behave differently from questions where hundreds of pages are equally plausible.
| Engine | Cluster | Answers | Domains | No vendor cited |
|---|---|---|---|---|
| openai | money-trial | 73 | 2 | 0% |
| openai | vs-comparison | 64 | 14 | 0% |
| openai | pricing-fees | 56 | 7 | 1.8% |
| openai | dev-services | 56 | 12 | 1.8% |
| openai | how-to | 48 | 3 | 0% |
| gemini | money-trial | 53 | 68 | 34% |
| gemini | vs-comparison | 26 | 203 | 38.5% |
| gemini | pricing-fees | 14 | 104 | 21.4% |
| gemini | dev-services | 14 | 137 | 57.1% |
| gemini | how-to | 1 | 12 | n too small |
Who actually gets cited
The open surface is not, as a rule, made of media brands. It is made of app vendors, theme shops and agencies publishing long explainer posts about the platform they build on.
gemini (gemini-2.5-flash)
Share of 108 answers
| Domain | Share | Med. rank |
|---|---|---|
| shopify.com | 63.9% | 6 |
| dodropshipping.com | 39.8% | 3 |
| pagefly.io | 38.9% | 2 |
| ecomposer.io | 28.7% | 5 |
| demandsage.com | 27.8% | 3 |
| shopify.ecom-store.pro | 26.9% | 5 |
| bsscommerce.com | 25% | 3 |
| litextension.com | 25% | 6 |
| hulkapps.com | 23.1% | 8 |
| folio3.com | 19.4% | 4 |
| reddit.com | 16.7% | 8 |
| stylefactoryproductions.com | 14.8% | 6 |
openai (gpt-5.4-mini)
Share of 297 answers
| Domain | Share | Med. rank |
|---|---|---|
| shopify.com | 96.3% | 1 |
| shopify.dev | 7.7% | 1 |
| woocommerce.com | 4.7% | 2 |
| wix.com | 3% | 3 |
| bigcommerce.com | 2% | 5 |
| google.com | 1.7% | 1 |
| squarespace.com | 1.7% | 3 |
| affirm.com | 1.3% | 2 |
| afterpay.com | 1.3% | 4 |
| klarna.com | 1.3% | 3 |
| printful.com | 1.3% | 2 |
| tiktok.com | 1.3% | 2 |
The right-hand column is the whole story of that engine: after the vendor and its developer documentation, the numbers collapse into the low single digits and stay there. The complete list of everything it cited in 297 answers is 32 domains long, and most of them are other platforms being compared against Shopify rather than publishers writing about it.
Where we land ourselves
Publishing a citation study without publishing your own line in it would be a strange thing to do, so: this site was cited in 29 of 405 answers. Every one of them came from gemini. On the openai leg the count is 0 out of 297, which is what the 80.1% figure above looks like from the receiving end.
When cited, the median position in the source list was 5, and 4 answers put this site first. The distribution across our own pages is lopsided in a way that turns out to be the most useful thing we learned:
| Page | Answers citing it |
|---|---|
| /blog/shopify-6-month-trial/ | 24 |
| /blog/shopify-vs-shop-app/ | 3 |
| /blog/add-shopify-to-existing-website/ | 1 |
| /blog/shopify-trial-billing-traps/ | 1 |
| /blog/shopify-1-dollar-promotion/ | 1 |
One page carries 24 of the 29 citations, and it is the page answering a dated, checkable question about a current promotion. In the clusters where the question has no single correct answer, our share is zero, and it is zero against a field of hundreds of pages that are all equally right. That is the pattern we would test next if we were starting this again: a citation appears to be much easier to win where the answer expires than where it is merely good.
Download the data
All three files are plain CSV, UTF-8, regenerated from the same parse that produced every figure above.
- ai-citation-panel-answers.csv
One row per answered prompt: date, engine, model, cluster, the prompt text itself, every domain cited in order, and whether this site was among them. 405 rows.
- ai-citation-domain-leaderboard.csv
Every domain either engine cited, with the number and share of answers citing it and its median position in the citation list.
- ai-citation-engine-cluster-summary.csv
The engine-by-cluster grid: answers, distinct domains, vendor share, publisher room, and this site's own share.
The runner that produces them is part of this site's open tooling; the measurement rules above are documented in its header rather than in a methodology note that can drift away from the code.
Reuse and attribution
The data and the charts are published under Creative Commons Attribution 4.0. Republish them anywhere, including commercially, including modified, as long as you credit the source with a link. No permission request needed and no reply to wait for.
Every chart is a standalone file:
- citation-surface-width.svg
- vendor-share-of-slot.svg
- publisher-room-by-cluster.svg
- top-cited-domains-gemini.svg
Each one carries its own source line inside the image, so the credit survives being downloaded and reuploaded elsewhere. To embed one with the attribution already in place:
<a href="https://shopify.ecom-store.pro/research/ai-citations-shopify/"><img src="https://shopify.ecom-store.pro/research/ai-citations-shopify/charts/citation-surface-width.svg" alt="Distinct domains cited in Shopify buyer-intent answers"></a>
<p>Source: <a href="https://shopify.ecom-store.pro/research/ai-citations-shopify/">Shopify Ecom AI Citation Panel</a>, CC BY 4.0</p>Preferred credit line: Shopify Ecom AI Citation Panel, with a link to this page. If you are writing something that needs a figure we did not publish, the raw file above almost certainly contains it.
About this measurement
Run by Alexander Matynian, a Shopify developer since 2017, on Shopify Ecom. The panel started as an internal instrument for tracking whether our own articles were being cited, which is why our own domain is one of the measured parties and why the raw file is published rather than summarised.
The written analysis on this site, including this page, is drafted with AI under human direction and verified before publication. See our editorial policy. The measurements themselves are machine-collected from the engines' citation metadata and are not AI-generated.
Questions or a correction: the data file is the fastest way to check anything here. Reach us through contacts.
