Same Buyer, Same Brand, Four Different Verdicts
I asked four AI engines whether a buyer should choose a brand. Then I asked again, 575 more times.
Earlier this month I came pretty close to publishing a study that was wrong. The numbers in it were real. I'd run 54 conversations where a simulated buyer talks to ChatGPT, Claude, Gemini, and Perplexity about a purchase, gets a short list, pushes back a few times, and finally asks, "So should I choose this brand?" I had seven challenger brands, a draft, charts, the whole thing. And the draft had some good lines in it. One of them was that every time an engine turned a challenger down, it sent the buyer to the category leader. Another was that Claude refuses to pick.
Then I went back and looked at how I'd run it, and I didn't love what I saw. Each brand and engine combination had been run one time. The buyer was an AI improvising its own wording, so no two conversations were the same. There were no category leaders in the test to compare against. And when I dug into the tool (it's my tool, I built it) I found that the buyer's script had been fed each brand's selling points, so some of those selling points could have leaked into what the buyer said.
So I rebuilt the test and ran it again, properly this time, and both of those good lines turned out to be wrong.
I've put the full study on its own page, with every chart, the method, the limits, and the data to download. This post is the shorter version, for people who mostly want to know what I found and what I'd do about it.
How the test works
I picked nine categories and, in each one, a challenger brand and the brand most people would call the leader. Pipedrive and HubSpot. Huntress and CrowdStrike. Omnisend and Klaviyo. Pixieset and Squarespace. Helix and Tempur-Pedic. And four more.
For each pair I wrote one buyer, and I wrote that buyer to be the kind of customer the challenger's own sales team would call a perfect fit. The CRM buyer, for example, is the founder of a 12-person commercial cleaning company with a budget under $50 per user and nobody on staff to administer software. I did that on purpose. A lot of people in marketing (me included, until I ran this) assume these systems default to the biggest name in the category because the biggest name has the most written about it. I wanted to see whether a perfectly matched small buyer still gets steered to the big brand.
The conversation is seven turns, and the buyer's side is a frozen script, identical word for word on every run. The first turn asks for options and doesn't name a brand. The last one asks "should I choose it?" The category leader gets the same seven messages with only the brand name swapped. Every conversation ran six times per engine over three days with web search on, and then again with search off. That's 576 conversations, and every final answer got labeled and then checked, including by two people working blind.
The leader got turned down about twice as often
Across all four engines, the category leader was rejected in 35% of conversations. The challenger was rejected in 16%. Challengers got a clean yes 57% of the time and leaders 26%, and those numbers barely moved from one day to the next.
What I saw was the engines reading the buyer's constraints and using them. HubSpot got rejected in 18 of 24 conversations, and the reasons were the ones I'd written into the buyer, the $50 cap and the lack of an admin. CrowdStrike got dinged for cost and complexity for a two-person IT team. And when an engine rejected a leader, it named the paired challenger in that same answer 55 times out of 75, without the buyer ever having mentioned it.
That takes care of "every dismissal goes to the category leader." With leaders in the test, the leaders got dismissed more than the challengers did. I'd have published the opposite of what the data says.
Two leaders never got a clean yes from any engine, 24 tries each. Tempur-Pedic was rejected outright by ChatGPT and Claude every time, mostly over heat and price. Squarespace is the one I'd look at if I ran marketing for a big general-purpose product. The buyer was a wedding photographer who wants client galleries and print sales, and 18 of the 24 answers were some version of what ChatGPT said, which was "Don't choose Squarespace as your only tool for client galleries and print sales," followed by advice to use it for the public site and add a specialist for the rest. That kind of answer doesn't show up as a loss in any dashboard I know of, and it's a loss.
Which engine you ask matters more than which day
If you only look at one chart, make it that grid, and look down the columns.
ChatGPT never rejected a challenger in 54 conversations. Gemini almost never rejects anybody, and it was the only engine that said yes to HubSpot, six times out of six. Perplexity sits in the middle. And then Claude (the Haiku model, which is the small one) rejected the challenger 31% of the time and the category leader 74% of the time.
Klaviyo is the cleanest example of what that means for a brand. It was the same buyer and the same seven messages over three days. ChatGPT said yes to Klaviyo six times out of six. Claude said no six times out of six. Neither one budged. If you're Klaviyo, what your "AI visibility" looks like depends on which of those two your buyer happens to have open, and there's nothing in a single blended score that would tell you that.
Same words in, different answer out
The buyer's messages never changed, and I can prove it, because the tool hashes them. So any difference between two runs of the same cell is the engine.
In 21 of 72 brand and engine pairings, which is 29%, the same buyer got a recommendation on some runs and a rejection on others. In 11 of them the swing was a clean yes on one run and a no on another. Omnisend on Claude went yes, yes, no, yes, no, yes. CrowdStrike on Perplexity went conditional, no, no, no, yes, no. ChatGPT and Gemini were steady, splitting on two brands each out of 18. Claude split on 11.
I think there are really two kinds of cell here. When an engine says no to you six times out of six, it has made up its mind, and that's a positioning problem or a product problem. When it says yes three times and no three times, it hasn't, and what it lands on probably depends on which pages it happened to pull that run. That second kind seems a lot more fixable to me.
What the engines believe when they can't look anything up
With web search off, ChatGPT and Gemini barely changed. Whatever those two think about Klaviyo or Brooks or HubSpot, they seem to already think it before they search.
Claude was a different animal. From memory, it said yes once in 36 conversations, and 11 times it declined to give a verdict at all, with lines like "I can't responsibly say yes or no." With search on, across 108 conversations, it never did that once. So "Claude refuses to pick," my other good line from the first draft, turned out to be true only when Claude can't search.
Huntress is the example that would worry me if I worked there. With search, Claude said yes six times out of six. From memory it said no both times, because "Huntress requires you to have security expertise to act on what it finds," which for a two-person IT team is disqualifying. With search, the same model described a managed service with a 24/7 human team. I'm not going to referee which description is right. I'd just want to know that one of the major models carries that belief around and only drops it when it goes and reads.
What I'd do with this
Number one, stop checking your brand once. I know that's annoying to hear, because running a prompt and screenshotting the answer is how most of us have been doing it, me included. However, in almost a third of these pairings a single check had a real chance of telling you the opposite of what usually happens. Run it several times, on different days, and look at the spread.
Two, check each engine separately and don't average them. An average of six yeses and six nos describes nothing.
Three, find out what the models believe about you with search turned off. If what's in there is out of date, the fix is a lot of consistent, plainly worded content in places other than your own site, over a long time, and I don't think there's a shortcut.
Four, say plainly who you're for. The engines quoted specifics back to the buyer, things like price per seat, team size, and what's included. That's how the challengers won, and it's also how the leaders lost the small buyer, so if you're the big brand it's worth deciding which of those buyers you want to fight for.
One caution. This study measured what the engines responded to. It didn't test whether changing any of it moves a verdict. That's the next one I want to run.
The full study has six interactive charts, the method, the limits (there are several, and I built the tool, so please do check my work), and every conversation in a downloadable file. It's all here. If you've run anything like this on your own brand, I'd love to see it, especially if it doesn't match mine.
Jarred Smith is the author of Explainable: Why AI Recommends Some Brands & Ignores Others, an Amazon bestseller on AEO, GEO, and SEO. He's a marketing leader with nearly 20 years of experience across healthcare, public media, retail, and environmental services. Find him at jarredsmith.com.