ChatGPT and Gemini disagree on the best software in a third of categories, and the review sites have been skipped

ChatGPT and Gemini name a different best product in a third of software categories, a study of 9,978 answers found. It also found the review sites have been skipped entirely.


ChatGPT and Gemini disagree on the best software in a third of categories, and the review sites have been skipped

GetIntel published its AI Software Index on Wednesday. The study asked roughly 80 buyer-phrased questions in each of 126 software categories, covering 1,825 brands. The questions went to the live ChatGPT and Gemini consumer apps rather than the developer APIs. That produced 9,978 usable answers, all collected on 6 August.

The headline result is a coin toss. In one category out of every three, the two assistants name a different number one. Which chatbot a buyer happens to open decides which company gets recommended.

The review sites have been skipped

Look at what the assistants actually cite. The most-cited domain across the whole study was techradar.com, with 1,114 citations. Reddit came second on 807 and Zapier third on 418. G2, Capterra and TrustRadius together accounted for 5% of all citations.

That is an industry being routed around. Those three sites exist to be the reference layer for software buying. A consumer tech magazine and a forum are now doing that job instead. Gartner managed 238 citations, fewer than Zapier.

Reddit sitting second is its own problem. The platform is already cracking down on brands seeding fake opinions for chatbots to repeat. It says it now catches 25,000 such posts a day. The assistants are leaning hardest on the source that is hardest to keep clean.

ChatGPT sends traffic. Gemini does not

The two behave very differently once they have picked a winner. When ChatGPT recommends a product, it cites that product’s own website 53% of the time. Gemini does so in only 13% of cases. It leans on third-party pages instead.

So the same recommendation is worth different amounts. Winning inside ChatGPT delivers a link. Winning inside Gemini delivers a mention. That distinction matters to anyone still measuring publisher traffic as the outcome.

Gemini’s independent sources are stranger still. Its most-cited third-party pages are blog posts written by other software companies. The study notes these are often about categories the publishing company has nothing to do with.

The answer changes with who is asking, and where

Saying who you are moves the result. In some categories a brand goes from winning 34% of answers to winning 90%. The only change is whether the buyer says “I’m a freelancer” or “we’re an enterprise”. Same category, same question, different customer.

Geography moves it too, and this is the finding a European reader should sit with. The number one recommendation flips between US-phrased and UK-phrased versions of the same question. That happened in about half the categories the study could measure. Regulators have spent two years arguing about whether AI answers opt out fairly. Nobody has asked whether they answer the same question the same way on both sides of the Atlantic.

Showing up and winning are not the same

The leaderboard data is unkind to the obvious names. HubSpot appears on more category leaderboards than anyone, 19 of them, and wins three. Salesforce appears on 13 and wins none at all. QuickBooks appears on seven and wins five.

Some categories are already settled. GitHub is the first recommendation in 91% of CI/CD answers and 72% of code review answers. Shopify takes 84% of e-commerce and Miro 83% of whiteboards. Square holds 78% of point of sale and Loom 76% of screen recording.

Others are wide open. In affiliate marketing software the top pick wins only 5% of answers. Contract management sits at 13%. Invoicing, live chat and order fulfilment all sit at 19%. The winner there changes almost every time the question is reworded.

Who measured this, and why

The caveat is the whole business model. GetIntel is a Bengaluru company that sells AI visibility software. Its full index concludes that companies do not know whether AI names them. That is also its sales pitch. Founder Tarang Agarwal says so nearly outright.

Agarwal puts the problem plainly. A company can quote its Google ranking to the decimal, he said. It may have no idea whether ChatGPT has ever named it. “Closing that gap is the entire reason an AI visibility tool exists.”

It is a crowded pitch. Berlin’s Peec AI reached a $200m valuation selling the same promise. The practical work of earning citations still looks a lot like ranking on Google.

The method has real limits and they are worth stating. This is one day of data from two engines, with no error bars and no peer review. “Top pick” simply means the first brand named in an answer, which is a blunt measure of a nuanced reply. A single snapshot cannot show whether any of this is stable.

Which is the test worth setting. GetIntel says it will refresh the Index quarterly. If the disagreement rate is still near a third in November, buyers really are getting different answers from different machines. If it collapses, this was a photograph of a moving object.

Get the TNW newsletter

Get the most important tech news in your inbox each week.