69.3% of ChatGPT's Software Recommendations Have No Source Behind Them
Across 60 software buying answers, ChatGPT made 358 brand recommendations. Only 110 of them were backed by a citation to that vendor anywhere in the answer. That is 248 unsupported recommendations, or 69.3%. Even the single most-recommended product in each answer – the "my pick would be" brand – was unsupported 28.3% of the time. The brand was asserted from model memory, not evidence.
The test I applied here is deliberately generous. A recommendation counts as supported if any citation in that answer points at the recommended vendor's own domain. It does not have to be the specific claim being made. It does not have to be a good source. It just has to exist. On that loose test, 248 of 358 recommendations fail.
The numbers
| Measure | Value |
|---|---|
| Brand recommendations in the corpus | 358 |
| Backed by a citation to that vendor | 110 |
| With no supporting citation | 248 (69.3%) |
| Unique brands recommended | 150 |
| Of top-ranked picks, share with no citation | 28.3% |
The pattern tightens at the top of the list. The headline pick in an answer is more likely than average to be sourced – but roughly one in four still rests on model memory alone.
What this looks like in a single answer
Aggregates hide mechanism. Two captures make the pattern concrete.
"What are the best alternatives to NetSuite?" ChatGPT read 34 results across nine domains, rendered seven citations from two domains – erpresearch.com five times and cudio.com twice – and named twelve brands: Oracle, NetSuite, Microsoft Dynamics 365 Business Central, Acumatica, SAP Business One, Odoo, Microsoft Dynamics 365, Sage X3, Epicor, IFS, Oracle Fusion and SAP S/4HANA. Zero of those twelve had a supporting citation to their own site. The entire shortlist rests on a single third-party comparison page plus prior belief.
"What is the best HR and payroll software?" 43 results read across seventeen domains. Two citations rendered, both to pricing pages – rippling.com and gusto.com – for the two brands whose prices ChatGPT happened to quote. ADP, Paychex and Deel were each positioned in the market with no supporting source at all. Answer latency: 4.0 seconds.
Notice what the citations are doing in the second example. They are not evidence for the recommendation. They are evidence for the price. The recommendation came first; the citation attached itself to the checkable fragment.
"What is practice management software in healthcare and who are the leading vendors in 2026?" The heaviest retrieval in the corpus: 120 results read across 46 domains, producing nine citations from five domains and fifteen named brands. Athenahealth was backed by a source. Tebra, AdvancedMD, NextGen, eClinicalWorks, Veradigm, ModMed, DrChrono, SimplePractice, CareCloud, Greenway Health, EMIS, TPP, SystmOne and Vision were not. Contested, fragmented categories drive enormous fan-out that the user never sees, and the extra reading barely changes the sourcing ratio.
Why this splits AI visibility into two problems
If 69.3% of recommendations arrive with no evidence attached, a large share of the outcome is decided before retrieval begins. That has a direct structural consequence for how you work on AI visibility.
Getting cited is a retrieval and content-structure problem. It is solvable with pages the fan-out can find: dated, titled, tabular, product-named, priced. It moves on a quarterly timescale, and it is measurable this week.
Getting recommended is a training-data and entity-salience problem. It is a question of whether the model already associates your brand with the category, at what rank, and in what framing. It moves on a much longer horizon and it shows up in the 28.3% of top picks that need no citation at all.
They are not the same game, and optimising for the first does not automatically win the second. A vendor can be cited constantly – because it has a great pricing page – and still lose the ranking. A vendor can be recommended first in every answer and never be cited once, because nothing about it was checkable.
Which brands the model already believes in
The most-recommended brands table is a rough proxy for entity salience in this corpus. "Answers naming it" is breadth; "times ranked #1" and "mean rank" are how strongly the model leads with it.
| Brand | Answers naming it | Times ranked #1 | Mean rank | Total mentions |
|---|---|---|---|---|
| HubSpot | 11 | 6 | 1.91 | 109 |
| Microsoft Dynamics 365 | 9 | 0 | 4.44 | 27 |
| QuickBooks | 6 | 3 | 1.83 | 32 |
| Xero | 6 | 2 | 1.67 | 58 |
| CrowdStrike | 6 | 3 | 1.83 | 51 |
| SentinelOne | 6 | 0 | 3.17 | 38 |
| Shopify | 6 | 3 | 1.50 | 80 |
| BigCommerce | 6 | 1 | 2.33 | 52 |
| Salesforce | 5 | 2 | 1.80 | 57 |
| Asana | 5 | 4 | 1.60 | 44 |
| Clio | 5 | 4 | 1.20 | 65 |
| MyCase | 5 | 1 | 1.80 | 42 |
| ClickUp | 5 | 1 | 2.80 | 28 |
| Monday.com | 5 | 0 | 2.80 | 29 |
| Zoho CRM | 5 | 0 | 3.60 | 29 |
| SAP S/4HANA | 5 | 2 | 4.40 | 15 |
| Pipedrive | 5 | 0 | 4.80 | 26 |
| NetSuite | 5 | 0 | 5.20 | 24 |
| Epicor | 5 | 1 | 5.80 | 26 |
| eClinicalWorks | 5 | 0 | 3.80 | 20 |
Two patterns worth flagging. Microsoft Dynamics 365 appears in nine answers – second only to HubSpot on breadth – and led none of them, with a mean rank of 4.44. It is the archetypal "always on the list, never the pick" position. Clio is the opposite: five answers, four of them as the top pick, mean rank 1.20. In legal practice management the model has an opinion and it is not close.
Neither of those positions is something a content programme changes this quarter.
What to actually do
Stop reporting "AI visibility" as one number. Split it. Citation share is one metric with one set of levers. Recommendation share and mean rank are a different metric with a different, slower set of levers. A dashboard that merges them will show you noise.
Make more of your answer checkable. Citations attach to verifiable fragments. If your pricing is gated, your feature matrix is a PDF and your integration list is a JavaScript-rendered carousel, you have removed the hooks that citations bind to – which is exactly what the accounting capture in this corpus shows, where one vendor with a clean public pricing page took all seven citations in the answer.
Treat mean rank as the entity metric. Times-ranked-first and mean position across a spread of prompt phrasings is a better proxy for model belief than raw mention count. HubSpot's 109 mentions and Xero's 58 tell you far less than HubSpot's 1.91 and Xero's 1.67.
Accept the timescale. Entity salience responds to sustained category presence – documentation, integrations ecosystem, editorial coverage, the whole surface that ends up in training data – not to a content sprint. Budget for it as a two-year programme, and measure the citation half quarterly while you wait.
FAQ
Are ChatGPT's software recommendations backed by sources?
Mostly not. Of 358 brand recommendations across 60 answers, 110 were backed by a citation to that vendor and 248 – 69.3% – were not. The test used was generous: any citation anywhere in the answer pointing at the recommended vendor's domain counted as support.
Is ChatGPT's top recommendation more likely to be sourced?
Yes, but not reliably. The single most-recommended product in each answer was unsupported 28.3% of the time, against 69.3% for recommendations overall. Roughly one headline pick in four rests entirely on model memory.
What is the difference between being cited and being recommended by AI?
Citation is a retrieval and content-structure outcome you can influence with findable, checkable pages on a quarterly timescale. Recommendation is an entity-salience outcome driven by what the model already associates with your category, and it moves far more slowly.
Why does ChatGPT recommend brands it has no source for?
Because the answer is composed from a blend of retrieved evidence and parametric memory. Brand rankings and "best for" judgements frequently come from the latter, while citations attach to the checkable fragments of the answer such as price points and feature availability.
How do I improve entity salience for AI recommendations?
Slowly and structurally. Sustained category presence - documentation, integrations, third-party coverage, consistent naming - is what feeds the training data the model reasons from. Track mean rank and times-ranked-first across a spread of prompt phrasings rather than raw mention counts.
Related in this series
- How ChatGPT Shortlists Software Brands, the full report as a PDF
- Forty-Five Reads, Six Links: How ChatGPT Actually Cites Software Brands
- Same Buyer, Same Category, Different Shortlist: Prompt Shape Rewrites the Answer
- Your AI Visibility Tool Is Blind to 79% of the Web ChatGPT Reads
- ERP in ChatGPT: Twelve Vendors Named, Zero Sources, One Comparison Site Deciding the Shortlist
About the research. Nathan Mzumara is an organic growth and AI search practitioner. Brand detection used a fixed 350-term lexicon scoped to each category, with overlapping matches masked longest-first so that "Microsoft Dynamics 365" is never double-counted as "Dynamics 365". Method and limitations are stated in the pillar report.
Tags