Do AI Assistants Agree on the Best Software? · What 80 B2B Categories Reveal
Cooper's August 16 snapshot shows unanimous Best leaders in only 17 of 80 B2B software categories, with much greater fragmentation around price, setup, and rising products.

Ask four AI assistants for the best software in a category and the result is rarely one shared answer. In Cooper's August 16, 2026 snapshot, ChatGPT, Gemini, Claude, and Perplexity named the same #1 product in only 17 of 80 B2B software categories. Among the 66 Best boards where all four assistants supplied a usable #1, 49 produced some disagreement.
That is the central finding of this study: AI software recommendations do not form one market. They form overlapping recommendation markets, each shaped by a different assistant, question, retrieval path, and response. Consensus is meaningful when it appears, but disagreement is the normal condition.
For buyers, the practical response is to compare assistants and then verify the shortlist against requirements and primary evidence. For software companies, it means that one blended visibility score cannot explain who is recommending the product, for which buyer question, or how durable that position may be.
The Short Answer
- Unanimous Best leaders were the exception: all four assistants chose the same #1 in 17 of 80 categories.
- Most complete Best boards still had a majority: at least three assistants agreed on the #1 in 39 of the 66 categories with four usable assistant results.
- Constraint questions fragmented much more: only 2 complete category boards were unanimous on Most Affordable, and only 2 were unanimous on Rising.
- Three Best markets split four ways: CPQ, Procurement, and Cloud Hosting (PaaS) had a different #1 from every assistant.
- Agreement is not proof of product quality: it describes recommendation convergence in a dated observation, not a universal verdict.
What Cooper Measured
This analysis uses Cooper's immutable August 16, 2026 ranking snapshot. It covers 80 B2B software categories across Sales, Marketing, Support, HR, Finance, Productivity, IT & Security, and Engineering. For each category, Cooper records ranked lists from ChatGPT, Gemini, Claude, and Perplexity for four direct buying questions:
- Best
- Most Affordable
- Easiest to Set Up
- Rising
For this article, agreement means two assistants placed the same Cooper product identity at #1 for the same category and question. Product identities remain separate where Cooper's published data treats them as separate. For example, a product family and one of its named products are not merged merely because they share a company.
A complete board is a category-question combination with a usable #1 from all four assistants. Missing results remain missing observations. They are not converted into zero ranks and do not count as disagreement in the distribution of one, two, three, or four distinct leaders.
The study compares the raw assistant #1 positions first. Cooper's published AIQ ranking is a separate category-relative calculation with the weights explained in the Cooper methodology. That distinction matters because equal editorial treatment of the four assistants does not mean the AIQ formula weights them equally.
Consensus Is the Exception, Not the Baseline
The broad “Best” question produced the most agreement, but even there unanimity was limited. Seventeen categories had one shared #1. Thirty had two distinct leaders, 16 had three, and 3 had four different leaders.
| Distinct #1 Products on Complete Best Boards | Categories | Share of 66 Complete Boards |
|---|---|---|
| 1, unanimous | 17 | 25.8% |
| 2 | 30 | 45.5% |
| 3 | 16 | 24.2% |
| 4, every assistant differed | 3 | 4.5% |
The 17 unanimous results are 21.3% of all 80 categories and 25.8% of the 66 complete boards. Both denominators are useful. The first describes the entire published category universe. The second isolates the boards where a four-way comparison is possible.
There is also more middle-ground consensus than the unanimity figure alone suggests. In 39 of 66 complete Best boards, at least three assistants chose the same leader. Cooper's weighted category leader was named #1 by at least three assistants in 41 of all 80 categories, including complete and incomplete boards. That is evidence of meaningful overlap, but it still leaves nearly half the category set without that level of convergence.
Buyer Constraints Change the Recommendation Market
The biggest shift appears when the question moves from a broad “Best” to a specific buying constraint. Most Affordable, Easiest to Set Up, and Rising produce substantially more fragmented #1 results.
| Buying Question | Complete Boards | Unanimous #1 | Four Different #1s |
|---|---|---|---|
| Best | 66 of 80 | 17 | 3 |
| Most Affordable | 68 of 80 | 2 | 27 |
| Easiest to Set Up | 70 of 80 | 9 | 22 |
| Rising | 72 of 80 | 2 | 32 |
The pattern is not surprising, but its size matters. “Best” can reward established category leaders with broad recognition. “Affordable” depends on what price, plan, company size, and total cost the assistant implicitly treats as relevant. “Easy” can mean a short setup, a familiar interface, limited administration, or strong implementation support. “Rising” requires a judgment about momentum, which is especially sensitive to recent evidence and category interpretation.
This is why a software company can lead a generic recommendation query and disappear when the buyer adds a real constraint. It is also why visibility should be reported by buying question. An average across all four questions would hide the difference between broad category authority and a specific product-market fit.
Where All Four Assistants Agree
The unanimous Best leaders span all eight Cooper hubs, with Engineering contributing three categories. Each of these products reached AIQ 100 in its category because all four assistants placed it first.
Unanimity is a strong visibility signal. It means the product won the same direct question across four separate recommendation surfaces in this snapshot. It does not prove that the product is objectively best for every buyer, nor does it explain why the assistants agreed.
The evidence footprint should still be read separately. For example, the complete Conversation Intelligence category contained 111 Sources and 448 Mentions across all direct questions, while Proposal Software contained 148 Sources and 571 Mentions. Those are category-wide evidence measurements, not product-specific explanations for the ranks of
Gong or PandaDoc. Cooper defines Sources as distinct cited pages and Mentions as verified product-page occurrences across those cited pages.
Three Markets Where Every Assistant Picks a Different Leader
At the other end of the spectrum, three Best categories produced four different #1 products.
| Hub | Category | ChatGPT | Gemini | Claude | Perplexity | Sources | Mentions |
|---|---|---|---|---|---|---|---|
| Sales | CPQ | 147 | 566 | ||||
| Finance | Procurement | 115 | 251 | ||||
| Engineering | Cloud Hosting (PaaS) | 126 | 467 |
These are not low-evidence categories. CPQ had 147 distinct cited pages and 566 verified Mentions across the category's completed direct questions. Cloud Hosting had 126 Sources and 467 Mentions. The disagreement therefore cannot be reduced to a simple absence of retrievable material.
The more useful interpretation is that the category leaves room for different evaluation frames. In CPQ, even product identity is consequential: ChatGPT and Gemini selected two distinct Salesforce products, while Claude and Perplexity chose different vendors. In Cloud Hosting, the four leaders represent different positions across developer experience, cloud breadth, deployment workflow, and platform maturity. A short “best” prompt does not tell the assistant which of those criteria should dominate.
Cooper's weighted Best leader was Salesforce Revenue Cloud in CPQ at AIQ 58,
SAP Ariba in Procurement at AIQ 55, and
Render in Cloud Hosting at AIQ 75. Those rankings answer Cooper's published weighted comparison. They do not erase the underlying four-way split.
Pairwise Agreement Is Uneven
The Best results also show that assistant pairs do not agree at one common rate. Claude shared the same #1 with ChatGPT in 48 of 80 categories, the highest pairwise result in this snapshot. ChatGPT and Gemini agreed in 30 of the 66 categories where both supplied a usable Best leader.
| Assistant Pair | Same Best #1 | Comparable Categories | Agreement Rate |
|---|---|---|---|
| ChatGPT and Claude | 48 | 80 | 60.0% |
| Claude and Perplexity | 44 | 80 | 55.0% |
| Gemini and Claude | 33 | 66 | 50.0% |
| ChatGPT and Perplexity | 39 | 80 | 48.8% |
| Gemini and Perplexity | 32 | 66 | 48.5% |
| ChatGPT and Gemini | 30 | 66 | 45.5% |
The denominators differ because Gemini did not supply a usable Best #1 in 14 categories in the published snapshot. That absence is itself operationally relevant, but it should not be treated as either agreement or disagreement.
These rates describe this dated category set, not a permanent relationship between assistants. A different prompt library, date, account state, market, language, or product universe can produce a different pattern.
Why AI Assistants Can Disagree
No provider publishes a complete commercial recommendation formula, so the observed differences should not be reverse-engineered into unsupported ranking-factor claims. The providers do document enough of their search behavior to show why one answer pipeline should not be assumed to match another.
- OpenAI explains that ChatGPT Search may rewrite a question into one or more targeted queries, consult search providers, and issue narrower follow-up queries after reviewing initial results.
- Google documents that Gemini's search-grounding workflow analyzes the prompt, decides whether search can improve the answer, generates one or multiple queries, processes the results, and synthesizes a cited response.
- Anthropic states that Claude invokes a search tool for current information, processes multiple sources, and returns citations and source links.
- Perplexity describes a real-time internet search process that interprets the question, gathers information, summarizes it, and includes numbered citations.
Those descriptions do not reveal the assistants' product-ranking logic. They do establish that each surface performs its own prompt interpretation, retrieval, processing, and synthesis. Different results can emerge before model preference even enters the picture.
The wording of the question is another source of variation. A 2026 Web Science paper, Self-Promotion in LLM Recommendations, used three paraphrases of each base question because semantically similar prompts can produce different recommendations. The researchers also queried each model five times per prompt because outputs are stochastic. Their subject was recommendations for AI models, not B2B software, so their commercial-bias findings should not be generalized to Cooper's categories. Their audit design does support a broader measurement lesson: standardize the question, repeat the observation, and preserve the provider-level result.
What Agreement Does and Does Not Prove
Agreement proves something narrow and valuable: multiple assistants independently surfaced the same product at the same position for the same defined question in a recorded snapshot.
It does not prove:
- that the product is the best fit for a particular company's requirements;
- that the assistants used the same sources or evaluation criteria;
- that the cited pages caused the recommendation;
- that the result will persist after a model, index, interface, or market change;
- that an assistant's fluent explanation is factually complete.
OpenAI's own search guidance warns that results and citations can be “incomplete, outdated, or incorrect” and tells users to open the cited source and verify that it supports the answer. That is a sensible standard across assistants.
A 2024 experiment with 199 participants, published in Frontiers in Computer Science, found no significant difference in trust or reliance between recommendations presented as LLM-generated and AI-sourced snippets. Trust fell after inaccurate advice. The study used three general-knowledge tasks and pre-generated recommendations, so it does not measure B2B software choice. It does reinforce the practical importance of accuracy feedback and calibrated trust.
What Software Buyers Should Do
Treat an assistant recommendation as the beginning of a shortlist, not the end of procurement.
- Ask the same buying question in more than one assistant. The disagreement is informative. It reveals where a shortlist is stable and where the market depends on the surface.
- Add the constraints that actually decide the purchase. Company size, region, budget, implementation resources, integrations, data residency, and security requirements can change the answer materially.
- Inspect why each product was recommended. Separate claims about fit, price, ease, momentum, and category leadership instead of accepting “best” as one dimension.
- Open the cited pages. Confirm that they exist, are current, and support the adjacent claim. Prefer vendor documentation for product facts and independent evidence for comparative claims.
- Build a human-verified shortlist. Use demos, references, security reviews, contract terms, and implementation evidence before making a decision.
Cross-assistant consensus can raise confidence that a product has broad recommendation visibility. It should never replace buyer-specific due diligence.
What Software Companies Should Measure
The study points to a more useful AI visibility operating model.
- Report each assistant separately. A blended score is a summary, not a diagnosis.
- Segment by buyer question. Best, affordable, easy, and rising are different recommendation markets.
- Track agreement and coverage. A #1 position in one assistant is different from a majority or unanimous lead.
- Keep rank separate from evidence. Sources are distinct cited pages; Mentions are verified product-page occurrences. Neither metric is recommendation rank.
- Repeat strategically. One snapshot establishes what happened on a date. Repeated runs establish whether the position is stable.
- Preserve the exact observation context. Record date, assistant surface, question, category, product identity rules, failures, and methodology changes.
Our guide to measuring AI search visibility provides a full prompt, sampling, evidence, and reporting framework. The companion research on which sources shape software recommendations shows how cited-page patterns can be studied without confusing evidence volume with rank.
Limitations
This is an observational analysis of one Cooper snapshot dated August 16, 2026. It compares the published #1 product from each assistant and question, not every rank in every list. Fourteen Best boards lacked a usable Gemini #1, with smaller missing sets in the other questions. Those results were excluded from complete-board distributions and pairwise comparisons where necessary.
The article does not test causality, recommendation accuracy, buyer outcomes, prompt paraphrases, geography, account personalization, or repeated responses from the same assistant. It does not infer undisclosed ranking factors from the cited pages. The category taxonomy and product identities follow Cooper's published artifact.
The results should therefore be read as a transparent baseline: broad enough to reveal a strong cross-category pattern, but dated and specific enough to be reproduced and challenged.
Frequently Asked Questions
Do ChatGPT, Gemini, Claude, and Perplexity Recommend the Same Software?
Sometimes, but not usually at #1. In Cooper's August 16, 2026 snapshot, all four assistants chose the same Best leader in 17 of 80 B2B software categories. Among the 66 categories with four usable Best leaders, 49 contained some disagreement.
Which Buying Question Creates the Most Agreement?
Best produced the most unanimous leaders in this snapshot, with 17 complete category boards. Easiest to Set Up had 9, while Most Affordable and Rising had 2 each.
Does Four-Assistant Agreement Mean a Product Is Objectively Best?
No. It means four assistants placed the same product first for a defined question in a dated observation. A buyer still needs to verify requirements, product facts, comparative claims, implementation fit, and commercial terms.
Why Are Some Comparison Denominators Smaller Than 80?
Some published category-question boards did not contain a usable #1 from every assistant. Cooper treats those as missing observations, not zero ranks. Complete-board and pairwise rates use only the categories where the required assistant results exist.
Should AI Visibility Be One Overall Score?
An overall score can help scan a market if its weighting is explicit. It should remain traceable to assistant-level ranks, buying questions, Sources, Mentions, and dated observations. Otherwise it conceals the disagreement that explains the result.
The Practical Standard
AI assistants agree often enough to create recognizable category leaders, but not often enough to behave like one recommendation engine. The clearest signal in Cooper's 80-category snapshot is the coexistence of consensus and fragmentation: 17 unanimous Best leaders, 30 two-leader markets, 16 three-leader markets, and 3 complete four-way splits.
Buyers should compare, constrain, and verify. Software teams should measure by assistant and question before rolling anything into a composite. The market is not one answer, and the measurement system should not pretend that it is.
Related News

How to Measure AI Search Visibility · Metrics That Make the Data Defensible
A practical measurement system for prompt design, repeated sampling, assistant-level reporting, evidence analysis, and choosing an AI visibility tool without trusting a black-box score.

Where AI Software Recommendations Find Their Evidence This Week
The current edition adds 173 Sources and 1,693 Mentions overall, but eight hub examples show that evidence growth remains uneven at category level.

23 of 80 Best Leaders Changed in This Week’s Software Rankings
This week's edition maps 2,853 ranked products across 80 categories, with 9,119 Sources, 36,148 Mentions, and a clearer divide between settled leaders and active markets.