Changing question phrasing dramatically alters AI product recommendations in software leaderboards

A groundbreaking experiment reveals that how questions are posed to AI models can have a greater impact on recommended products than the models themselves, reshaping how market visibility is measured.

A detailed experiment in AI-driven software discovery suggests that the wording of a question can matter more than the model answering it. The author behind connexion.me said he published free leaderboards tracking which products ChatGPT, Gemini and Perplexity name when prompted to recommend software, and then tested the same buying intent in 44 different phrasings across two CRM categories. The headline result was stark: small-business CRM and open-source CRM boards produced no overlap in their top 20 products, despite appearing to cover what might seem like the same market.

The method was deliberately repetitive. Each question was asked in 44 variations, every answer was counted and the raw runs were published so the results could be checked. According to the write-up, changing the engine while holding the wording steady left most of the top 10 intact, with at least 7 of 10 products remaining in common across ChatGPT, Gemini and Perplexity. But when the wording changed and the engine stayed fixed, the overlap collapsed to almost nothing.

The clearest example came from SuiteCRM. On the open-source board, it appeared in all 44 phrasings and in 81 of 88 answers for that category, yet it did not appear at all in 132 answers to the small-business CRM prompts. The author said a forum reply from the SuiteCRM community argued that the original questions lacked enough operational detail, and that adding constraints such as sensitive data or on-premises needs would likely surface different results.

He then tested that criticism by rewriting each prompt with more business context while keeping the underlying buying question unchanged. In that follow-up run, SuiteCRM moved from absent to dominant: the post says it was named 42 times out of 44 by ChatGPT, 33 times by Gemini and 8 times by Perplexity. The broader board also shifted substantially, with top-10 overlap between the original and context-rich versions ranging from 4 of 10 to 6 of 10 depending on the engine.

That led to the post’s main conclusion: these leaderboards do not simply measure a market, they measure a phrasing family. Industry tools that audit AI visibility are built around the same basic idea. AgentGEO, for example, offers brand audits across ChatGPT, Perplexity and Gemini, while Orbator’s recommendation index tracks which CRM tools AI assistants mention most often and refreshes weekly. TechRadar has also described answer engine optimisation as an emerging discipline that depends on identifying prompts, measuring mentions and improving source signals.

The author was careful not to overstate the findings. He noted that the extra context run was not yet published alongside the live boards, that the sample covered one intent rather than a full market and that a name being mentioned is not the same as being recommended. Even so, the exercise suggests that AI answers can swing sharply when the buyer’s situation is made more specific, which has obvious implications for vendors trying to understand how their products surface in AI-generated comparisons.

Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.