Same Buyer, Same Category, Different Shortlist: Prompt Shape Rewrites the Answer
Six different phrasings of the same buying question returned brand sets that overlapped by a mean pairwise Jaccard similarity of only 0.338. Across ten software categories the top-ranked vendor changed 3.8 times on average out of six prompts. In HR and payroll it changed all six times: Rippling, Paylocity, Workday, ADP, BrightHR and Oracle HCM each led exactly once. This is the largest single source of variance in the corpus, and it means any single tracked prompt is close to noise.
I held the category constant and varied only the shape of the buying question. Ten categories, six archetypes each, sixty prompts. The archetypes are the six ways a real buyer asks about software.
The six archetypes
| # | Archetype | Example prompt |
|---|---|---|
| A | Unconstrained best-of | What is the best CRM software? |
| B | Segment-constrained best-of | What is the best CRM software for a 200-employee UK company? |
| C | Head-to-head comparison | Compare Salesforce vs HubSpot for a mid-market B2B sales team. |
| D | Alternatives-to | What are the best alternatives to Salesforce? |
| E | Budget / requirement constrained | I need a CRM under $50 per user per month that integrates with Xero and Outlook. |
| F | Category + vendor landscape | What is CRM software and who are the leading vendors in 2026? |
Same category. Same buyer. Six answers that behave like six different products.
What changes when the shape changes
| # | Archetype | Results read | Citations | Cited domains | Brands named | Words | Vendor-owned citation share |
|---|---|---|---|---|---|---|---|
| A | Unconstrained best-of | 43.2 | 4.6 | 3.3 | 6.3 | 404 | 47.8% |
| B | Segment-constrained best-of | 41.8 | 6.0 | 3.4 | 5.1 | 827 | 68.3% |
| C | Head-to-head comparison | 41.3 | 6.7 | 2.7 | 2.1 | 1,034 | 91.0% |
| D | Alternatives-to | 37.7 | 6.9 | 4.2 | 7.8 | 741 | 47.8% |
| E | Budget / requirement constrained | 52.3 | 6.5 | 3.8 | 3.6 | 460 | 92.3% |
| F | Category + vendor landscape | 51.7 | 6.0 | 4.1 | 10.9 | 929 | 41.7% |
Two columns move in opposition, and the mechanism is clean.
A head-to-head comparison names 2.1 brands and draws 91.0% of its citations from the two vendors' own sites. The question has already picked the shortlist, so retrieval is spent verifying feature and price claims rather than shopping. When the buyer supplies the names, ChatGPT stops shopping and starts fact-checking – which makes the vendor's own site the entire evidence base. The CrowdStrike versus SentinelOne capture in this corpus read 43 results from just three domains and sent all eight citations to the two competitors' own product and pricing pages.
A budget-constrained question does the same thing for a different reason: 92.3% vendor citations, because only the vendor publishes the price. The accounting capture is the purest case – 27 results read from five domains, seven citations, all seven from zoho.com, the single vendor recommended.
At the other end, a category-landscape question names 10.9 brands and takes only 41.7% of citations from vendors. That is the archetype where analysts, aggregators and comparison sites actually get a hearing. It is also the second-longest answer at 929 words, behind head-to-head at 1,034, and one of the heaviest retrievals at 51.7 results.
If you are a challenger brand, archetypes D and F are your entry points. If you are the incumbent being compared, archetype C is where you win or lose on your own pages.
How unstable is the shortlist?
For each category I compared the brand sets returned by the six archetypes pairwise, using Jaccard similarity. A score of 1.0 would mean phrasing makes no difference. The corpus mean is 0.338 – barely a third of the brands overlap between two ways of asking the same question.
| Category | Shortlist stability (mean pairwise Jaccard) |
|---|---|
| CRM | 0.529 |
| Cybersecurity | 0.430 |
| Accounting | 0.426 |
| ERP / manufacturing | 0.373 |
| Ecommerce platforms | 0.355 |
| Project management | 0.336 |
| Legal practice management | 0.287 |
| Healthcare EHR | 0.247 |
| Marketing automation | 0.216 |
| HR & payroll | 0.185 |
CRM is the most settled category at 0.529. HR and payroll is near-chaotic at 0.185.
The leader changes 3.8 times out of six
Stability of the set is one thing. Stability of the winner is what most vendors actually care about.
| Category | Distinct leaders in 6 prompts | Most frequent leader | Its share |
|---|---|---|---|
| HR & payroll | 6 | Rippling | 17% |
| ERP / manufacturing | 5 | SAP S/4HANA | 33% |
| Marketing automation | 5 | HubSpot | 33% |
| Cybersecurity | 4 | CrowdStrike | 50% |
| Ecommerce platforms | 4 | Shopify | 50% |
| Accounting | 3 | QuickBooks | 50% |
| Healthcare EHR | 3 | Athenahealth | 50% |
| Legal practice management | 3 | Clio | 67% |
| Project management | 3 | Asana | 67% |
| CRM | 2 | HubSpot | 67% |
In HR and payroll the top-ranked vendor was different in all six prompts. There is no "ChatGPT thinks the best HR software is X". There is only what it says to a particular question.
Even in the most settled category in the corpus, CRM, the leader changed twice out of six and the most frequent leader held only 67% of the top slots.
What this breaks, and what to do instead
It breaks single-prompt tracking. If your AI visibility tool reports "you rank #2 for 'best HR software'", that is one observation from a distribution with six distinct outcomes. Track it weekly and you will see movement that is phrasing artefact, not model drift.
It breaks single-number reporting. "Our AI visibility is 34%" is meaningless without knowing which archetypes are in the denominator and how much they vary. A brand strong in head-to-head and absent from category landscape has a completely different problem from the reverse, and both average out to the same headline number.
It breaks naive win/loss narratives. A vendor that "lost the top spot in ChatGPT this month" may simply have been measured on a different phrasing.
The alternative is straightforward, if less comfortable to put on a slide.
Measure the prompt space, not the prompt. Run all six archetypes per category. Report the distribution – mean rank, times ranked first, presence rate – with its variance stated alongside. A brand at mean rank 2.1 with low variance is in a materially better position than one at mean rank 1.8 with high variance, and a single-prompt tracker cannot tell you which you are.
Segment your archetype performance. Your citation share in archetype C tells you whether your own pages are doing their job. Your presence rate in archetype F tells you whether third parties list you. Those are separate diagnoses with separate fixes.
Set the sample size honestly. With a mean overlap of 0.338, six observations per category is a probe, not a census. Category-level figures here should be read as directional. If you are running this internally on your own category, more phrasings and repeated runs will tighten the estimate considerably.
Watch the archetype mix in your category. Categories with high vendor-citation share concentrate in archetypes C and E, where your own site is the evidence base. Categories with low vendor share concentrate in F, where you need third-party listings. Knowing which you are in determines where the work goes.
FAQ
How much does prompt phrasing change ChatGPT's software shortlist?
Substantially. Across ten categories, six phrasings of the same buying question returned brand sets overlapping by a mean pairwise Jaccard similarity of 0.338. Roughly two thirds of the brands differ between two ways of asking the same thing.
Does ChatGPT have a consistent "best" software for a category?
Usually not. The top-ranked vendor changed 3.8 times on average across six prompts per category. In HR and payroll it changed all six times, with Rippling, Paylocity, Workday, ADP, BrightHR and Oracle HCM each leading exactly once.
Which prompt types favour vendor websites?
Head-to-head comparisons at 91.0% vendor citation share and budget-constrained questions at 92.3%. Both narrow the answer to checkable claims about named products, where the vendor's own page is the primary source. Category-landscape questions are the opposite at 41.7%.
Which prompt type names the most brands?
Category and vendor landscape questions, at a mean of 10.9 brands named per answer, followed by alternatives-to at 7.8. Head-to-head comparisons name the fewest at 2.1, because the question has already chosen the shortlist.
How should AI visibility be tracked given this variance?
As a distribution across prompt archetypes, not a single tracked prompt. Report mean rank, times ranked first and presence rate with variance stated. Segment results by archetype, since strength in head-to-head and absence from category landscape are different problems with different fixes.
Related in this series
- How ChatGPT Shortlists Software Brands, the full report as a PDF
- Forty-Five Reads, Six Links: How ChatGPT Actually Cites Software Brands
- Your AI Visibility Tool Is Blind to 79% of the Web ChatGPT Reads
- HR Software in ChatGPT: Six Prompts, Six Different Winners
- Cybersecurity in ChatGPT: When the Buyer Names the Vendors, Your Own Site Becomes the Whole Evidence Base
About the research. Nathan Mzumara is an organic growth and AI search practitioner. Each prompt ran in a fresh conversation, 54 of 60 in temporary chat. Brand sets were compared pairwise using Jaccard similarity on normalised brand names. Method and limitations are stated in the pillar report.
Tags