All insights
AI search research
8 min read26 August 2026Nathan Mzumara

Vendor Pages Take 65.9% of ChatGPT's Citations. G2 Takes 4.1%.

Vendor Pages Take 65.9% of ChatGPT's Citations. G2 Takes 4.1%.

Across 367 citations rendered in 60 software buying answers, 65.9% pointed at a software vendor's own website. Independent comparison sites took 14.4%. The five established review aggregators – G2, Capterra, Software Advice, TrustRadius and TechnologyAdvice – took 4.1% between them. Community and user-generated content took 0.0%: Reddit was read six times, LinkedIn six times and YouTube nineteen times, and not one of those retrievals ever became a citation.

The mental model most marketers carry into AI search is that answers are assembled from third-party authority. Reviews. Analyst notes. Community consensus. That model is intuitive, it matches how buyers behave, and in this corpus it is wrong.

What ChatGPT reads versus what ChatGPT cites

I hand-classified all 225 cited-or-repeatedly-retrieved domains into ten source types. Where a domain's nature was not obvious, I read the site before classifying it. Here is the full picture – share of the 2,680 retrievals against share of the 367 citations. The final row is the residual: domains retrieved fewer than twice and never cited.

Source typeDomainsRetrievalsRetrieval shareCitationsCitation shareConversion
Vendor-owned791,47655.1%24265.9%16.4%
Independent comparison site432499.3%5314.4%21.3%
Agency / consultancy / partner582067.7%215.7%10.2%
Analyst / market research10873.2%164.4%18.4%
Editorial media20813.0%154.1%18.5%
Review aggregator5521.9%154.1%28.8%
Independent testing lab2220.8%30.8%13.6%
Government / regulator4170.6%20.5%11.8%
Community / UGC311.2%00.0%0.0%
Reference160.2%00.0%0.0%
Long-tail (uncited, residual)33945316.9%00.0%0.0%

Two columns matter more than the rest. The gap between retrieval share and citation share tells you which source types get promoted when the model decides what to link. The conversion column tells you the odds that any given retrieval from that source type survives.

Vendor pages are not a fallback. They are the primary source.

Vendor-owned pages take 65.9% of citations from 55.1% of retrievals. That is a promotion, not a default.

This is not the model deferring to marketing copy for opinions. It is the model sourcing facts. Price points, tier names, feature availability, integration support, supported platforms. For all of those, the vendor's own page genuinely is the primary source, and the model treats it that way. Pricing pages alone accounted for 278 of the 2,680 retrieved results.

The concentration is starker than the average suggests. In 26 of 60 answers – 43% – every single citation was a vendor page. No third party appeared at all. The buyer sees a confident, specific, well-sourced answer in which the only evidence is the seller's own website.

The clearest case in the corpus was a budget-constrained accounting question: "I need accounting software under $40 a month that handles multi-currency invoicing and US sales tax." ChatGPT read 27 results from five domains, rendered seven citations, and all seven came from zoho.com – the single vendor it recommended. Xero and QuickBooks were both named in the answer with no supporting citation at all.

Review aggregators are trusted but barely surfaced

The most counter-intuitive number in this dataset is the aggregator conversion rate: 28.8%, the highest of any source type.

When ChatGPT retrieves a G2 or Capterra page, it is more likely to cite it than a page from any other category. The aggregators have not lost the trust battle. They have lost the retrieval battle. The whole category contributed just 52 retrievals across sixty answers – 1.9% of everything read – and 15 citations.

Being trusted is worthless if the query fan-out never surfaces you. That is a very different diagnosis from "AI doesn't value reviews", and it points at a very different fix.

Community content converts at zero

Community and user-generated content is the inverse pattern: surfaced occasionally, cited never. 31 retrievals across sixty answers, zero citations. Reddit read six times. LinkedIn six times. YouTube nineteen times. Not one survived into a rendered answer.

This should be read carefully. It does not mean community discussion has no influence – it may well shape the model's parametric priors over long training horizons, which is exactly where the unsourced 69.3% of brand recommendations comes from. It means community content is not the evidence layer for purchase-intent software answers. If your AI visibility programme is built around seeding Reddit threads, this corpus offers no support for the citation half of that thesis.

The 20 most-cited domains

118 distinct domains earned at least one citation. Here are the twenty that earned six or more.

Vendor Pages Take 65.9% of ChatGPT's Citations. G2 Takes 4.1%.
DomainCitationsAnswersSource type
hubspot.com187Vendor
shopify.com155Vendor
xero.com145Vendor
zoho.com125Vendor
erpresearch.com114Independent comparison site
ciopages.com104Independent comparison site
salesforce.com96Vendor
microsoft.com94Vendor
asana.com94Vendor
clickup.com93Vendor
crowdstrike.com94Vendor
mycase.com82Vendor
adobe.com72Vendor
clio.com7Vendor
sage.com64Vendor
gartner.com64Analyst
g2.com63Review aggregator
monday.com63Vendor
bigcommerce.com63Vendor
tebra.com63Vendor

Sixteen of the top twenty are vendors. Two are small independent comparison sites most brand-monitoring tools have never heard of. One is Gartner. G2 appears at six citations across three answers – respectable, and roughly one third of what HubSpot's own domain achieved on its own.

Citation is extremely fragmented

There is no small set of publishers to court. The corpus has a Herfindahl–Hirschman index of just 177 across cited domains.

Concentration measureValue
Distinct domains cited118
Herfindahl–Hirschman index177
Top domain share4.9%
Top 5 share19.1%
Top 10 share31.6%
Top 20 share49.9%
Domains cited exactly once55 (46.6%)

The single most-cited domain in the entire corpus holds 4.9% of citations. Nearly half of all cited domains were cited exactly once. For comparison, an HHI of 177 would be considered an unconcentrated market by any competition authority in the world.

The strategic reading: there is no publisher relations play that moves this meaningfully. There is a very long tail to be present in, and one property you fully control that is worth more than the rest combined.

What to do with this

Treat your own domain as the primary citation asset. 65.9% of citations, and the only source in 43% of answers. The pages that earn this are unglamorous: pricing, plan comparison, feature availability, integration lists, supported-platform tables. Anything that states a checkable fact in a structured, scannable form.

Publish your prices. This is the single strongest lever in the dataset and it has its own article in this series. Categories with public pricing had vendor citation shares in the 80–91% range; categories with gated pricing sat at 48–55%, with the difference going to analysts and comparison sites.

Stop treating G2 rank as your AI visibility metric. Fifteen of 367 citations. Track it if you like, but it is not the mechanism.

Do not conclude that community is worthless. Conclude that community is not the citation layer. It may still be doing work in the 69.3% of recommendations that arrive with no evidence attached at all.

FAQ

Does ChatGPT cite G2 and Capterra?

Rarely. In this corpus of 367 citations, G2, Capterra, Software Advice, TrustRadius and TechnologyAdvice supplied 15 citations between them – 4.1% of the total. When one is retrieved it converts well, at 28.8%, the best of any source type, but the whole category was retrieved only 52 times across sixty answers.

Does ChatGPT cite Reddit for software recommendations?

Not in this corpus. Reddit was retrieved six times, LinkedIn six times and YouTube nineteen times across sixty purchase-intent software answers. None of those 31 community retrievals became a citation. Community content converted at 0.0%.

What share of ChatGPT citations go to vendor websites?

65.9%. Vendor-owned pages supplied 242 of 367 citations while accounting for 55.1% of retrievals, so they are promoted at citation time rather than merely being read most. In 26 of 60 answers, every citation in the answer was a vendor page.

Why does ChatGPT prefer vendor pages over review sites?

Because most cited claims are checkable facts rather than opinions. Source pills bind to pricing, tier names, feature availability and integration support, and the vendor's own page is the primary source for those. Pricing pages alone accounted for 278 of 2,680 retrieved results.

How concentrated are ChatGPT's citations?

Barely at all. 118 distinct domains were cited, with a Herfindahl–Hirschman index of 177. The top domain held 4.9% of citations, the top ten held 31.6%, and 55 domains – 46.6% of those cited – were cited exactly once.

About the research. Nathan Mzumara is an organic growth and AI search practitioner. All figures are computed from a live capture of the ChatGPT response stream on 26 August 2026: 60 prompts, 2,680 retrieved results, 367 citations, 225 hand-classified domains. Method and limitations are stated in the pillar report.

Tags

GEOAEOAI visibilitycitationsChatGPT

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

Downloading How ChatGPT Shortlists Software Brands also subscribes you to the Discovery Digest, one email every Friday. No spam, unsubscribe anytime.