All insights
AI search research
7 min read26 August 2026Nathan Mzumara

69.3% of ChatGPT's Software Recommendations Have No Source Behind Them

69.3% of ChatGPT's Software Recommendations Have No Source Behind Them

Across 60 software buying answers, ChatGPT made 358 brand recommendations. Only 110 of them were backed by a citation to that vendor anywhere in the answer. That is 248 unsupported recommendations, or 69.3%. Even the single most-recommended product in each answer – the "my pick would be" brand – was unsupported 28.3% of the time. The brand was asserted from model memory, not evidence.

The test I applied here is deliberately generous. A recommendation counts as supported if any citation in that answer points at the recommended vendor's own domain. It does not have to be the specific claim being made. It does not have to be a good source. It just has to exist. On that loose test, 248 of 358 recommendations fail.

The numbers

MeasureValue
Brand recommendations in the corpus358
Backed by a citation to that vendor110
With no supporting citation248 (69.3%)
Unique brands recommended150
Of top-ranked picks, share with no citation28.3%

The pattern tightens at the top of the list. The headline pick in an answer is more likely than average to be sourced – but roughly one in four still rests on model memory alone.

What this looks like in a single answer

Aggregates hide mechanism. Two captures make the pattern concrete.

"What are the best alternatives to NetSuite?" ChatGPT read 34 results across nine domains, rendered seven citations from two domains – erpresearch.com five times and cudio.com twice – and named twelve brands: Oracle, NetSuite, Microsoft Dynamics 365 Business Central, Acumatica, SAP Business One, Odoo, Microsoft Dynamics 365, Sage X3, Epicor, IFS, Oracle Fusion and SAP S/4HANA. Zero of those twelve had a supporting citation to their own site. The entire shortlist rests on a single third-party comparison page plus prior belief.

"What is the best HR and payroll software?" 43 results read across seventeen domains. Two citations rendered, both to pricing pages – rippling.com and gusto.com – for the two brands whose prices ChatGPT happened to quote. ADP, Paychex and Deel were each positioned in the market with no supporting source at all. Answer latency: 4.0 seconds.

Notice what the citations are doing in the second example. They are not evidence for the recommendation. They are evidence for the price. The recommendation came first; the citation attached itself to the checkable fragment.

"What is practice management software in healthcare and who are the leading vendors in 2026?" The heaviest retrieval in the corpus: 120 results read across 46 domains, producing nine citations from five domains and fifteen named brands. Athenahealth was backed by a source. Tebra, AdvancedMD, NextGen, eClinicalWorks, Veradigm, ModMed, DrChrono, SimplePractice, CareCloud, Greenway Health, EMIS, TPP, SystmOne and Vision were not. Contested, fragmented categories drive enormous fan-out that the user never sees, and the extra reading barely changes the sourcing ratio.

Why this splits AI visibility into two problems

If 69.3% of recommendations arrive with no evidence attached, a large share of the outcome is decided before retrieval begins. That has a direct structural consequence for how you work on AI visibility.

Getting cited is a retrieval and content-structure problem. It is solvable with pages the fan-out can find: dated, titled, tabular, product-named, priced. It moves on a quarterly timescale, and it is measurable this week.

Getting recommended is a training-data and entity-salience problem. It is a question of whether the model already associates your brand with the category, at what rank, and in what framing. It moves on a much longer horizon and it shows up in the 28.3% of top picks that need no citation at all.

They are not the same game, and optimising for the first does not automatically win the second. A vendor can be cited constantly – because it has a great pricing page – and still lose the ranking. A vendor can be recommended first in every answer and never be cited once, because nothing about it was checkable.

Which brands the model already believes in

The most-recommended brands table is a rough proxy for entity salience in this corpus. "Answers naming it" is breadth; "times ranked #1" and "mean rank" are how strongly the model leads with it.

69.3% of ChatGPT's Software Recommendations Have No Source Behind Them
BrandAnswers naming itTimes ranked #1Mean rankTotal mentions
HubSpot1161.91109
Microsoft Dynamics 365904.4427
QuickBooks631.8332
Xero621.6758
CrowdStrike631.8351
SentinelOne603.1738
Shopify631.5080
BigCommerce612.3352
Salesforce521.8057
Asana541.6044
Clio541.2065
MyCase511.8042
ClickUp512.8028
Monday.com502.8029
Zoho CRM503.6029
SAP S/4HANA524.4015
Pipedrive504.8026
NetSuite505.2024
Epicor515.8026
eClinicalWorks503.8020

Two patterns worth flagging. Microsoft Dynamics 365 appears in nine answers – second only to HubSpot on breadth – and led none of them, with a mean rank of 4.44. It is the archetypal "always on the list, never the pick" position. Clio is the opposite: five answers, four of them as the top pick, mean rank 1.20. In legal practice management the model has an opinion and it is not close.

Neither of those positions is something a content programme changes this quarter.

What to actually do

Stop reporting "AI visibility" as one number. Split it. Citation share is one metric with one set of levers. Recommendation share and mean rank are a different metric with a different, slower set of levers. A dashboard that merges them will show you noise.

Make more of your answer checkable. Citations attach to verifiable fragments. If your pricing is gated, your feature matrix is a PDF and your integration list is a JavaScript-rendered carousel, you have removed the hooks that citations bind to – which is exactly what the accounting capture in this corpus shows, where one vendor with a clean public pricing page took all seven citations in the answer.

Treat mean rank as the entity metric. Times-ranked-first and mean position across a spread of prompt phrasings is a better proxy for model belief than raw mention count. HubSpot's 109 mentions and Xero's 58 tell you far less than HubSpot's 1.91 and Xero's 1.67.

Accept the timescale. Entity salience responds to sustained category presence – documentation, integrations ecosystem, editorial coverage, the whole surface that ends up in training data – not to a content sprint. Budget for it as a two-year programme, and measure the citation half quarterly while you wait.

FAQ

Are ChatGPT's software recommendations backed by sources?

Mostly not. Of 358 brand recommendations across 60 answers, 110 were backed by a citation to that vendor and 248 – 69.3% – were not. The test used was generous: any citation anywhere in the answer pointing at the recommended vendor's domain counted as support.

Is ChatGPT's top recommendation more likely to be sourced?

Yes, but not reliably. The single most-recommended product in each answer was unsupported 28.3% of the time, against 69.3% for recommendations overall. Roughly one headline pick in four rests entirely on model memory.

What is the difference between being cited and being recommended by AI?

Citation is a retrieval and content-structure outcome you can influence with findable, checkable pages on a quarterly timescale. Recommendation is an entity-salience outcome driven by what the model already associates with your category, and it moves far more slowly.

Why does ChatGPT recommend brands it has no source for?

Because the answer is composed from a blend of retrieved evidence and parametric memory. Brand rankings and "best for" judgements frequently come from the latter, while citations attach to the checkable fragments of the answer such as price points and feature availability.

How do I improve entity salience for AI recommendations?

Slowly and structurally. Sustained category presence - documentation, integrations, third-party coverage, consistent naming - is what feeds the training data the model reasons from. Track mean rank and times-ranked-first across a spread of prompt phrasings rather than raw mention counts.

About the research. Nathan Mzumara is an organic growth and AI search practitioner. Brand detection used a fixed 350-term lexicon scoped to each category, with overlapping matches masked longest-first so that "Microsoft Dynamics 365" is never double-counted as "Dynamics 365". Method and limitations are stated in the pillar report.

Tags

GEOAEOAI visibilityentity salienceChatGPT

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

Downloading How ChatGPT Shortlists Software Brands also subscribes you to the Discovery Digest, one email every Friday. No spam, unsubscribe anytime.