The AI Verdict Study

What four AI engines told the same buyer about 18 brands, asked 576 times.

I wanted to know what happens after the first answer. A buyer asks an AI engine for options, gets a list, pushes back a few times, and then asks the question that matters, which is "so should I choose this one?" This study measures how that last answer comes out, for a challenger brand and for the category leader it competes with, across four engines, with the buyer's words held identical on every run.

If you want the story of how this came together (including the first version I nearly published, which was wrong), that's in the blog post. This page is the reference that includes the findings, the charts, the method, the limits, and the data.

Findings

  1. Category leaders were rejected about twice as often as challengers. Engines told the buyer not to choose the category leader in 35% of conversations (75 of 215) and the challenger in 16% (35 of 214). Challengers got a clean yes 57% of the time and leaders 26%. In every pair the buyer was written as the challenger's ideal customer, so this is a finding about fit. The engines read the buyer's constraints and used them.

  2. The engine mattered more than anything else measured. ChatGPT never rejected a challenger in 54 conversations. Claude rejected the category leader 74% of the time. Gemini rejected almost nobody (12% and 11%). Klaviyo got six yeses out of six from ChatGPT and six nos out of six from Claude, from the same buyer using the same words.

  3. 29% of brand and engine pairings were unstable. In 21 of 72 pairings, the same buyer got a recommendation on some runs and a rejection on others. In 11 of those, the answer went from a clean yes to a no. ChatGPT and Gemini each split on 2 brands of 18, Perplexity on 6, and Claude on 11.

  4. Being named in the first answer went with a better ending. When the engine named the brand in turn one, before the buyer mentioned it, the conversation ended in a rejection 18% of the time. When it didn't, 44%.

  5. Turning search off barely changed ChatGPT and Gemini, and changed Claude a lot. From memory alone, ChatGPT and Gemini matched their usual search-on verdict in 29 and 30 of 36 conversations. Claude matched in 17, said yes once, and refused to give any verdict 11 times. With search on, Claude never refused in 108 conversations.

  6. Two category leaders never got a clean yes. Squarespace (0 of 24) was usually recast as half the answer, with a specialist gallery tool recommended for the rest. Tempur-Pedic (0 of 24) was named in the first answer once and rejected outright by ChatGPT and Claude every time.

Challenger versus leader, by engine

The category leader got turned down twice as often

Share of conversations that ended with the engine telling the buyer not to choose the brand. In every pair, the buyer was written to be the challenger's ideal customer, and the leader got the identical buyer.

Rejected means a final answer of no, or a refusal to answer. Clean yes means the engine told the buyer to choose the brand without hanging the answer on something the buyer hadn't said. Two human raters agreed on rejected versus not rejected 90% of the time (kappa 0.78), and on the full four labels 66% of the time (kappa 0.53), so rejected is the sturdier measure. 95% intervals, all engines: challengers rejected 12 to 22%, leaders 29 to 41%. Source: AI Verdict Study v1, jarredsmith.com. 429 search-on conversations, Sept 18 to 20, 2026; 108 search-off conversations, Sept 20. gpt-5, claude-haiku-4-5, gemini-2.5-flash, sonar-pro.

The gap held on each of the three days. Challengers were rejected 18%, 14%, and 17% of the time, and leaders 35%, 36%, and 34%. When an engine rejected a leader, it named the paired challenger in that same answer 55 times out of 75, though the buyer had never mentioned it.

Counts are out of 24 conversations per brand (23 where a Gemini run didn't finish). Six challengers beat their leader. Brooks and Breville held. Basecamp and Asana both lost the agency buyer.

Every verdict, every run

Same buyer, same question, asked six times

Each dot is one full seven-turn conversation that ended with the buyer asking "should I choose this brand?" The buyer's words were identical every time. Tap a cell to read how each run ended.

Select a cell to see what the engine said in each run.
YesConditionalNoRefused to answer

With search: six runs per cell across three days. From memory: two runs per cell; Perplexity cannot turn search off. Three Gemini runs did not finish. Source: AI Verdict Study v1, jarredsmith.com. 429 search-on conversations, Sept 18 to 20, 2026; 108 search-off conversations, Sept 20. gpt-5, claude-haiku-4-5, gemini-2.5-flash, sonar-pro.

Each dot is one full conversation. Switch to the memory-only view to see what changed with web search off. Tap any cell to read how each run ended, in the engine's own words.

Stability

How often the answer depended on the day you asked

Each engine answered for 18 brands, six times each, with identical buyer wording. A brand counts as split when the same engine recommended it in some runs and rejected it in others. Tap a striped bar to see which brands.

Never rejected in six runsSplit: recommended and rejectedRejected all six runs
Across all four engines, 21 of 72 brand and engine pairings were split. In 11 of them the swing ran from a clean yes to a no.

Runs were spread over three consecutive days, two per day. Within a single day the two runs usually agreed; most of the movement happened between days. Source: AI Verdict Study v1, jarredsmith.com. 429 search-on conversations, Sept 18 to 20, 2026; 108 search-off conversations, Sept 20. gpt-5, claude-haiku-4-5, gemini-2.5-flash, sonar-pro.

YesConditionalNoRefused to answer

The buyer's messages are hashed, and the hash matched on every run, so any difference between two runs of the same cell comes from the engine. Two runs on the same day landed on the same side (favorable or rejected) in 183 of 213 same-day pairs. More of the movement showed up between days. As an out-of-sample check, a fourth day of Perplexity runs matched the most common verdict from the first three days 21 times out of 36.

First mention and final verdict

Brands named in the first answer were rejected less than half as often

Turn one never mentions a brand. The buyer describes the situation and asks for the best options. The brand only comes up in turn two, when the buyer says "I've heard of it." This compares how conversations ended depending on whether the engine had already named the brand unprompted.

How often each engine named the brand unprompted

Search-on conversations only. Final verdicts: 301 conversations where the brand was named in turn one, 128 where it wasn't. Mention rates include the three unfinished runs (432 total). ChatGPT often answered the shoe and mattress buyers with clarifying questions and no brands, which pulls its leader figure down. Source: AI Verdict Study v1, jarredsmith.com. 429 search-on conversations, Sept 18 to 20, 2026; 108 search-off conversations, Sept 20. gpt-5, claude-haiku-4-5, gemini-2.5-flash, sonar-pro.

YesConditionalRejected

Turn one never names a brand. This is a correlation, and part of it is a fit signal, since an engine that thinks a brand suits the buyer will tend to both name it early and recommend it late. ChatGPT's figure for leaders is pulled down by the shoe and mattress buyers, where it often answered turn one with clarifying questions and no brands.

Memory versus search

Only Claude changed when search was turned off

The same conversations, run once with web search on and once with it off, so the engine could only answer from what it learned in training. Perplexity is left out because its search can't be turned off.

YesConditionalNoRefused to answer

With search: 54 conversations per engine per group (52 and 53 for Gemini). From memory: 18 per engine per group, all on Sept 20, so treat the memory figures as directional. Claude refused to give any verdict in 11 of its 36 memory-only conversations and in none of its 108 with search. Source: AI Verdict Study v1, jarredsmith.com. 429 search-on conversations, Sept 18 to 20, 2026; 108 search-off conversations, Sept 20. gpt-5, claude-haiku-4-5, gemini-2.5-flash, sonar-pro.

This is the smallest sample in the study (two runs per cell, one day), so treat it as directional. Three cells show what search was doing for Claude. Huntress went from yes six times out of six with search to no both times from memory, where Claude's stated reason was "Huntress requires you to have security expertise to act on what it finds." With search, the same model described a managed service with a 24/7 human team. Saucony was told it "runs narrow" from memory and had its wide sizing cited with search. Pipedrive went the other way, a yes from memory and a no five times out of six with search, after Claude found cleaning-industry software (Jobber, QuoteIQ) and decided the buyer needed a field-service tool. I haven't tried to referee which descriptions are correct.

All 18 brands

All 18 brands, side by side

Every brand was judged 24 times: four engines, six runs each, web search on. Sort by any column. The pilot view shows the smaller first pass from Sept 16 and 17, which used a different method and is kept separate.

Yellow square: challenger brand. Blue square: category leader. Rejected means no or refused. Named first: share of conversations where the engine named the brand in turn one, before the buyer mentioned it. Source: AI Verdict Study v1, jarredsmith.com. 429 search-on conversations, Sept 18 to 20, 2026; 108 search-off conversations, Sept 20. gpt-5, claude-haiku-4-5, gemini-2.5-flash, sonar-pro.

Method

  • Pairs. Nine categories, each with a challenger and the brand most buyers would call the leader: Basecamp and Asana, Huntress and CrowdStrike, Omnisend and Klaviyo, Pipedrive and HubSpot, OnPay and Gusto, Pixieset and Squarespace, Saucony and Brooks, Helix and Tempur-Pedic, Gaggia and Breville.

  • Buyer. One buyer per pair, written as the challenger's ideal customer, with a role, a job to get done, and constraints stated as facts about the buyer (size, staffing, budget) and never as product features. Example: "founder and owner of a 12-person commercial cleaning company who handles sales personally," who needs to "move my 12-person commercial cleaning company off a lead-tracking spreadsheet and into a CRM," given "a budget under $50 per user per month and nobody on staff to administer software."

  • Script. Seven frozen turns: a broad question with no brand named, the brand introduced ("I've heard of..."), a skeptical question about tradeoffs given the constraints, a question about outgrowing it, a question about what people like me compare it with, a final pushback, and "should I choose it?" The leader's script is identical to the challenger's except for the brand name. Each brand's script is hashed with SHA-256, and the hash matched on every run, every day, and both search settings. The engine receives only the buyer's messages and its own earlier answers.

  • Engines. gpt-5 (returned as gpt-5-2025-08-07), claude-haiku-4-5 (20251001), gemini-2.5-flash, and Perplexity sonar-pro. All were called through their APIs by way of Replit AI Integrations (Perplexity through OpenRouter), with no system prompt, no date injected, and full history on every turn. Web search was each vendor's own tool. Claude was capped at five searches per turn because its API exposes a cap and the others don't. Temperature was left at provider defaults.

  • Runs. Search on: two runs per brand per engine per day on September 18, 19, and 20, 2026, for 432 conversations. Three Gemini runs on the third day didn't reach the final turn, so verdicts are out of 429. Search off: two runs per brand on ChatGPT, Claude, and Gemini on September 20, for 108. Thirty-six Perplexity runs launched in that batch are search-on by nature and were used only as the day-four check.

  • Labels. An automated classifier (gpt-5-mini) labeled each final answer yes, conditional, refused, or no, seeing only that answer and the brand name. A second reader screened all 573 final answers for labels that contradicted the answer's own wording and corrected 33. Two human raters then labeled 141 blind, with engine names hidden: all 33 corrected items plus a random 20% of the rest. On rejected (no or refused) versus not rejected, the raters agreed with each other 90% of the time (Cohen's kappa 0.78) and with the final labels 93% and 96% of the time. On the random sample, which stands in for the answers nobody hand-checked, agreement was 94% and 96%. On all four labels the raters agreed 66% of the time (kappa 0.53), with almost all disagreement falling between yes and conditional. So rejection rate is the primary measure and the clean-yes figures are secondary. Disagreements were settled by majority among the two raters and the second reader, and I broke one three-way tie. An answer of "use this brand for part of the job and something else for the rest" was labeled conditional.

  • Intervals. 95% Wilson intervals on the headline figures: challengers rejected 12 to 22%, leaders 29 to 41%; challengers clean yes 50 to 63%, leaders 21 to 32%.

Limits

  • One buyer per category, written to favor the challenger. I'm not claiming challengers beat leaders in general. A buyer written as the leader's ideal customer is a different study.

  • These are API results with no system prompt. The consumer apps add their own system prompts, memory, and search stack.

  • One model per vendor, and two of them (Haiku, Flash) are the small model in their family. I wouldn't assume the larger models behave the same way.

  • Three consecutive days. It says nothing about drift over weeks or months.

  • The buyer is a script. A real person would wander off it.

  • The memory-only round is two runs per cell on a single day.

  • The study measures what the engines responded to. It didn't test whether changing anything changes a verdict.

  • I built the tool that ran it (Citingly), so please do check my work.

The pilot

A first pass on September 16 and 17 ran seven of these challengers once per engine in each of two rounds, with an AI buyer that improvised its wording and no category leaders. It suggested two things that didn't survive the rerun which is that every dismissal went to the category leader, and that Claude refuses to pick. With leaders in the test, leaders were dismissed more often than challengers and Claude only refused when it couldn't search. The pilot's results are in the last chart as a separate tab and aren't combined with anything above.

Data

Every conversation, every turn, every label (the classifier's, the corrected one, and both raters'), and all 37,440 source URLs the engines cited are in one download.

Download the dataset (zip, 6 MB)

If you rerun any of this and get something different, I'd love to hear about it, especially if it disagrees with me.

How to cite

Smith, Jarred. "The AI Verdict Study" jarredsmith.com, September 2026. https://jarredsmith.com/research/ai-verdict-study

Changelog

  • Version 1, September 2026. Nine pairs, four engines, 432 search-on and 108 search-off conversations, audited labels with two-rater reliability check.

  • Planned. The same pairs with a buyer written as the leader's ideal customer. The larger model from each vendor. A before-and-after test of a specific content change on an unstable cell.

Jarred Smith is the author of Explainable: Why AI Recommends Some Brands & Ignores Others, an Amazon bestseller on AEO, GEO, and SEO. He's a marketing leader with nearly 20 years of experience across healthcare, public media, retail, and environmental services. Find him at jarredsmith.com.