Vendor Pages Take 65.9% of ChatGPT's Citations. G2 Takes 4.1%.
Across 367 citations rendered in 60 software buying answers, 65.9% pointed at a software vendor's own website. Independent comparison sites took 14.4%. The five established review aggregators – G2, Capterra, Software Advice, TrustRadius and TechnologyAdvice – took 4.1% between them. Community and user-generated content took 0.0%: Reddit was read six times, LinkedIn six times and YouTube nineteen times, and not one of those retrievals ever became a citation.
The mental model most marketers carry into AI search is that answers are assembled from third-party authority. Reviews. Analyst notes. Community consensus. That model is intuitive, it matches how buyers behave, and in this corpus it is wrong.
What ChatGPT reads versus what ChatGPT cites
I hand-classified all 225 cited-or-repeatedly-retrieved domains into ten source types. Where a domain's nature was not obvious, I read the site before classifying it. Here is the full picture – share of the 2,680 retrievals against share of the 367 citations. The final row is the residual: domains retrieved fewer than twice and never cited.
| Source type | Domains | Retrievals | Retrieval share | Citations | Citation share | Conversion |
|---|---|---|---|---|---|---|
| Vendor-owned | 79 | 1,476 | 55.1% | 242 | 65.9% | 16.4% |
| Independent comparison site | 43 | 249 | 9.3% | 53 | 14.4% | 21.3% |
| Agency / consultancy / partner | 58 | 206 | 7.7% | 21 | 5.7% | 10.2% |
| Analyst / market research | 10 | 87 | 3.2% | 16 | 4.4% | 18.4% |
| Editorial media | 20 | 81 | 3.0% | 15 | 4.1% | 18.5% |
| Review aggregator | 5 | 52 | 1.9% | 15 | 4.1% | 28.8% |
| Independent testing lab | 2 | 22 | 0.8% | 3 | 0.8% | 13.6% |
| Government / regulator | 4 | 17 | 0.6% | 2 | 0.5% | 11.8% |
| Community / UGC | – | 31 | 1.2% | 0 | 0.0% | 0.0% |
| Reference | 1 | 6 | 0.2% | 0 | 0.0% | 0.0% |
| Long-tail (uncited, residual) | 339 | 453 | 16.9% | 0 | 0.0% | 0.0% |
Two columns matter more than the rest. The gap between retrieval share and citation share tells you which source types get promoted when the model decides what to link. The conversion column tells you the odds that any given retrieval from that source type survives.
Vendor pages are not a fallback. They are the primary source.
Vendor-owned pages take 65.9% of citations from 55.1% of retrievals. That is a promotion, not a default.
This is not the model deferring to marketing copy for opinions. It is the model sourcing facts. Price points, tier names, feature availability, integration support, supported platforms. For all of those, the vendor's own page genuinely is the primary source, and the model treats it that way. Pricing pages alone accounted for 278 of the 2,680 retrieved results.
The concentration is starker than the average suggests. In 26 of 60 answers – 43% – every single citation was a vendor page. No third party appeared at all. The buyer sees a confident, specific, well-sourced answer in which the only evidence is the seller's own website.
The clearest case in the corpus was a budget-constrained accounting question: "I need accounting software under $40 a month that handles multi-currency invoicing and US sales tax." ChatGPT read 27 results from five domains, rendered seven citations, and all seven came from zoho.com – the single vendor it recommended. Xero and QuickBooks were both named in the answer with no supporting citation at all.
Review aggregators are trusted but barely surfaced
The most counter-intuitive number in this dataset is the aggregator conversion rate: 28.8%, the highest of any source type.
When ChatGPT retrieves a G2 or Capterra page, it is more likely to cite it than a page from any other category. The aggregators have not lost the trust battle. They have lost the retrieval battle. The whole category contributed just 52 retrievals across sixty answers – 1.9% of everything read – and 15 citations.
Being trusted is worthless if the query fan-out never surfaces you. That is a very different diagnosis from "AI doesn't value reviews", and it points at a very different fix.
Community content converts at zero
Community and user-generated content is the inverse pattern: surfaced occasionally, cited never. 31 retrievals across sixty answers, zero citations. Reddit read six times. LinkedIn six times. YouTube nineteen times. Not one survived into a rendered answer.
This should be read carefully. It does not mean community discussion has no influence – it may well shape the model's parametric priors over long training horizons, which is exactly where the unsourced 69.3% of brand recommendations comes from. It means community content is not the evidence layer for purchase-intent software answers. If your AI visibility programme is built around seeding Reddit threads, this corpus offers no support for the citation half of that thesis.
The 20 most-cited domains
118 distinct domains earned at least one citation. Here are the twenty that earned six or more.
| Domain | Citations | Answers | Source type |
|---|---|---|---|
| hubspot.com | 18 | 7 | Vendor |
| shopify.com | 15 | 5 | Vendor |
| xero.com | 14 | 5 | Vendor |
| zoho.com | 12 | 5 | Vendor |
| erpresearch.com | 11 | 4 | Independent comparison site |
| ciopages.com | 10 | 4 | Independent comparison site |
| salesforce.com | 9 | 6 | Vendor |
| microsoft.com | 9 | 4 | Vendor |
| asana.com | 9 | 4 | Vendor |
| clickup.com | 9 | 3 | Vendor |
| crowdstrike.com | 9 | 4 | Vendor |
| mycase.com | 8 | 2 | Vendor |
| adobe.com | 7 | 2 | Vendor |
| clio.com | 7 | – | Vendor |
| sage.com | 6 | 4 | Vendor |
| gartner.com | 6 | 4 | Analyst |
| g2.com | 6 | 3 | Review aggregator |
| monday.com | 6 | 3 | Vendor |
| bigcommerce.com | 6 | 3 | Vendor |
| tebra.com | 6 | 3 | Vendor |
Sixteen of the top twenty are vendors. Two are small independent comparison sites most brand-monitoring tools have never heard of. One is Gartner. G2 appears at six citations across three answers – respectable, and roughly one third of what HubSpot's own domain achieved on its own.
Citation is extremely fragmented
There is no small set of publishers to court. The corpus has a Herfindahl–Hirschman index of just 177 across cited domains.
| Concentration measure | Value |
|---|---|
| Distinct domains cited | 118 |
| Herfindahl–Hirschman index | 177 |
| Top domain share | 4.9% |
| Top 5 share | 19.1% |
| Top 10 share | 31.6% |
| Top 20 share | 49.9% |
| Domains cited exactly once | 55 (46.6%) |
The single most-cited domain in the entire corpus holds 4.9% of citations. Nearly half of all cited domains were cited exactly once. For comparison, an HHI of 177 would be considered an unconcentrated market by any competition authority in the world.
The strategic reading: there is no publisher relations play that moves this meaningfully. There is a very long tail to be present in, and one property you fully control that is worth more than the rest combined.
What to do with this
Treat your own domain as the primary citation asset. 65.9% of citations, and the only source in 43% of answers. The pages that earn this are unglamorous: pricing, plan comparison, feature availability, integration lists, supported-platform tables. Anything that states a checkable fact in a structured, scannable form.
Publish your prices. This is the single strongest lever in the dataset and it has its own article in this series. Categories with public pricing had vendor citation shares in the 80–91% range; categories with gated pricing sat at 48–55%, with the difference going to analysts and comparison sites.
Stop treating G2 rank as your AI visibility metric. Fifteen of 367 citations. Track it if you like, but it is not the mechanism.
Do not conclude that community is worthless. Conclude that community is not the citation layer. It may still be doing work in the 69.3% of recommendations that arrive with no evidence attached at all.
FAQ
Does ChatGPT cite G2 and Capterra?
Rarely. In this corpus of 367 citations, G2, Capterra, Software Advice, TrustRadius and TechnologyAdvice supplied 15 citations between them – 4.1% of the total. When one is retrieved it converts well, at 28.8%, the best of any source type, but the whole category was retrieved only 52 times across sixty answers.
Does ChatGPT cite Reddit for software recommendations?
Not in this corpus. Reddit was retrieved six times, LinkedIn six times and YouTube nineteen times across sixty purchase-intent software answers. None of those 31 community retrievals became a citation. Community content converted at 0.0%.
What share of ChatGPT citations go to vendor websites?
65.9%. Vendor-owned pages supplied 242 of 367 citations while accounting for 55.1% of retrievals, so they are promoted at citation time rather than merely being read most. In 26 of 60 answers, every citation in the answer was a vendor page.
Why does ChatGPT prefer vendor pages over review sites?
Because most cited claims are checkable facts rather than opinions. Source pills bind to pricing, tier names, feature availability and integration support, and the vendor's own page is the primary source for those. Pricing pages alone accounted for 278 of 2,680 retrieved results.
How concentrated are ChatGPT's citations?
Barely at all. 118 distinct domains were cited, with a Herfindahl–Hirschman index of 177. The top domain held 4.9% of citations, the top ten held 31.6%, and 55 domains – 46.6% of those cited – were cited exactly once.
Related in this series
- How ChatGPT Shortlists Software Brands, the full report as a PDF
- Forty-Five Reads, Six Links: How ChatGPT Actually Cites Software Brands
- The Sites Beating G2 in ChatGPT: 43 Comparison Sites, 53 Citations, 3.5x the Aggregators
- Publish Your Prices: Why Gated Pricing Costs You 30 Points of AI Citation Share
- The Page Shapes AI Retrieval Actually Finds: 525 Versus Pages, 17 Reviews
About the research. Nathan Mzumara is an organic growth and AI search practitioner. All figures are computed from a live capture of the ChatGPT response stream on 26 August 2026: 60 prompts, 2,680 retrieved results, 367 citations, 225 hand-classified domains. Method and limitations are stated in the pillar report.
Tags