All insights
AI search research
7 min read26 August 2026Nathan Mzumara

Same Buyer, Same Category, Different Shortlist: Prompt Shape Rewrites the Answer

Same Buyer, Same Category, Different Shortlist: Prompt Shape Rewrites the Answer

Six different phrasings of the same buying question returned brand sets that overlapped by a mean pairwise Jaccard similarity of only 0.338. Across ten software categories the top-ranked vendor changed 3.8 times on average out of six prompts. In HR and payroll it changed all six times: Rippling, Paylocity, Workday, ADP, BrightHR and Oracle HCM each led exactly once. This is the largest single source of variance in the corpus, and it means any single tracked prompt is close to noise.

I held the category constant and varied only the shape of the buying question. Ten categories, six archetypes each, sixty prompts. The archetypes are the six ways a real buyer asks about software.

The six archetypes

#ArchetypeExample prompt
AUnconstrained best-ofWhat is the best CRM software?
BSegment-constrained best-ofWhat is the best CRM software for a 200-employee UK company?
CHead-to-head comparisonCompare Salesforce vs HubSpot for a mid-market B2B sales team.
DAlternatives-toWhat are the best alternatives to Salesforce?
EBudget / requirement constrainedI need a CRM under $50 per user per month that integrates with Xero and Outlook.
FCategory + vendor landscapeWhat is CRM software and who are the leading vendors in 2026?

Same category. Same buyer. Six answers that behave like six different products.

What changes when the shape changes

#ArchetypeResults readCitationsCited domainsBrands namedWordsVendor-owned citation share
AUnconstrained best-of43.24.63.36.340447.8%
BSegment-constrained best-of41.86.03.45.182768.3%
CHead-to-head comparison41.36.72.72.11,03491.0%
DAlternatives-to37.76.94.27.874147.8%
EBudget / requirement constrained52.36.53.83.646092.3%
FCategory + vendor landscape51.76.04.110.992941.7%

Two columns move in opposition, and the mechanism is clean.

A head-to-head comparison names 2.1 brands and draws 91.0% of its citations from the two vendors' own sites. The question has already picked the shortlist, so retrieval is spent verifying feature and price claims rather than shopping. When the buyer supplies the names, ChatGPT stops shopping and starts fact-checking – which makes the vendor's own site the entire evidence base. The CrowdStrike versus SentinelOne capture in this corpus read 43 results from just three domains and sent all eight citations to the two competitors' own product and pricing pages.

A budget-constrained question does the same thing for a different reason: 92.3% vendor citations, because only the vendor publishes the price. The accounting capture is the purest case – 27 results read from five domains, seven citations, all seven from zoho.com, the single vendor recommended.

At the other end, a category-landscape question names 10.9 brands and takes only 41.7% of citations from vendors. That is the archetype where analysts, aggregators and comparison sites actually get a hearing. It is also the second-longest answer at 929 words, behind head-to-head at 1,034, and one of the heaviest retrievals at 51.7 results.

If you are a challenger brand, archetypes D and F are your entry points. If you are the incumbent being compared, archetype C is where you win or lose on your own pages.

How unstable is the shortlist?

For each category I compared the brand sets returned by the six archetypes pairwise, using Jaccard similarity. A score of 1.0 would mean phrasing makes no difference. The corpus mean is 0.338 – barely a third of the brands overlap between two ways of asking the same question.

CategoryShortlist stability (mean pairwise Jaccard)
CRM0.529
Cybersecurity0.430
Accounting0.426
ERP / manufacturing0.373
Ecommerce platforms0.355
Project management0.336
Legal practice management0.287
Healthcare EHR0.247
Marketing automation0.216
HR & payroll0.185

CRM is the most settled category at 0.529. HR and payroll is near-chaotic at 0.185.

The leader changes 3.8 times out of six

Stability of the set is one thing. Stability of the winner is what most vendors actually care about.

CategoryDistinct leaders in 6 promptsMost frequent leaderIts share
HR & payroll6Rippling17%
ERP / manufacturing5SAP S/4HANA33%
Marketing automation5HubSpot33%
Cybersecurity4CrowdStrike50%
Ecommerce platforms4Shopify50%
Accounting3QuickBooks50%
Healthcare EHR3Athenahealth50%
Legal practice management3Clio67%
Project management3Asana67%
CRM2HubSpot67%

In HR and payroll the top-ranked vendor was different in all six prompts. There is no "ChatGPT thinks the best HR software is X". There is only what it says to a particular question.

Even in the most settled category in the corpus, CRM, the leader changed twice out of six and the most frequent leader held only 67% of the top slots.

What this breaks, and what to do instead

It breaks single-prompt tracking. If your AI visibility tool reports "you rank #2 for 'best HR software'", that is one observation from a distribution with six distinct outcomes. Track it weekly and you will see movement that is phrasing artefact, not model drift.

Same Buyer, Same Category, Different Shortlist: Prompt Shape Rewrites the Answer

It breaks single-number reporting. "Our AI visibility is 34%" is meaningless without knowing which archetypes are in the denominator and how much they vary. A brand strong in head-to-head and absent from category landscape has a completely different problem from the reverse, and both average out to the same headline number.

It breaks naive win/loss narratives. A vendor that "lost the top spot in ChatGPT this month" may simply have been measured on a different phrasing.

The alternative is straightforward, if less comfortable to put on a slide.

Measure the prompt space, not the prompt. Run all six archetypes per category. Report the distribution – mean rank, times ranked first, presence rate – with its variance stated alongside. A brand at mean rank 2.1 with low variance is in a materially better position than one at mean rank 1.8 with high variance, and a single-prompt tracker cannot tell you which you are.

Segment your archetype performance. Your citation share in archetype C tells you whether your own pages are doing their job. Your presence rate in archetype F tells you whether third parties list you. Those are separate diagnoses with separate fixes.

Set the sample size honestly. With a mean overlap of 0.338, six observations per category is a probe, not a census. Category-level figures here should be read as directional. If you are running this internally on your own category, more phrasings and repeated runs will tighten the estimate considerably.

Watch the archetype mix in your category. Categories with high vendor-citation share concentrate in archetypes C and E, where your own site is the evidence base. Categories with low vendor share concentrate in F, where you need third-party listings. Knowing which you are in determines where the work goes.

FAQ

How much does prompt phrasing change ChatGPT's software shortlist?

Substantially. Across ten categories, six phrasings of the same buying question returned brand sets overlapping by a mean pairwise Jaccard similarity of 0.338. Roughly two thirds of the brands differ between two ways of asking the same thing.

Does ChatGPT have a consistent "best" software for a category?

Usually not. The top-ranked vendor changed 3.8 times on average across six prompts per category. In HR and payroll it changed all six times, with Rippling, Paylocity, Workday, ADP, BrightHR and Oracle HCM each leading exactly once.

Which prompt types favour vendor websites?

Head-to-head comparisons at 91.0% vendor citation share and budget-constrained questions at 92.3%. Both narrow the answer to checkable claims about named products, where the vendor's own page is the primary source. Category-landscape questions are the opposite at 41.7%.

Which prompt type names the most brands?

Category and vendor landscape questions, at a mean of 10.9 brands named per answer, followed by alternatives-to at 7.8. Head-to-head comparisons name the fewest at 2.1, because the question has already chosen the shortlist.

How should AI visibility be tracked given this variance?

As a distribution across prompt archetypes, not a single tracked prompt. Report mean rank, times ranked first and presence rate with variance stated. Segment results by archetype, since strength in head-to-head and absence from category landscape are different problems with different fixes.

About the research. Nathan Mzumara is an organic growth and AI search practitioner. Each prompt ran in a fresh conversation, 54 of 60 in temporary chat. Brand sets were compared pairwise using Jaccard similarity on normalised brand names. Method and limitations are stated in the pillar report.

Tags

GEOAEOAI visibilitymeasurementChatGPT

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

Downloading How ChatGPT Shortlists Software Brands also subscribes you to the Discovery Digest, one email every Friday. No spam, unsubscribe anytime.