Getting Shortlisted by AI Is Stable. Your Rank Isn't.

Jim Wrubel

Jim Wrubel

10/13/2026

#AI Search Visibility#prompt tracking#AI recommendations#ChatGPT#Gemini#Claude#research
Getting Shortlisted by AI Is Stable. Your Rank Isn't.

Ask an AI assistant the same buying question twice and you'll often get two different answers. The vendors shuffle. The one named first on Monday is third on Tuesday. If you track AI visibility, you've seen this, and you've probably had to explain it to a client or a boss who watched their brand "drop."

So which part of an AI answer can you count on? Our research team ran a pre-registered study to find out. We asked ChatGPT, Gemini and Claude the same 40 B2B software buying questions every day for a week and checked which vendors each answer recommended. The part that holds is the shortlist. When ChatGPT recommended a vendor for a question, it recommended that vendor again 81% of the time when we asked the same question on another day. Claude was close behind. Gemini was less consistent. The vendor's exact position moved far more.

Chasing rank is futile. Chasing a place on the shortlist isn't. The full study has the method and the data. Here's what it means for you.

Key takeaways

  • When ChatGPT or Claude recommends a B2B software vendor for a buying question, it usually recommends that vendor again when the same question is asked on another day.
  • Gemini's shortlist changes more from day to day than ChatGPT's or Claude's, so a single Gemini answer tells you less.
  • A vendor's exact position in an answer moves a lot from run to run. Whether it makes the shortlist mostly doesn't.
  • The top pick rotates, but usually among vendors that stay on the shortlist. Losing the top spot rarely means losing the recommendation.
  • What AI says a vendor is "best for" changes more than whether AI recommends the vendor at all.
  • ChatGPT, Gemini and Claude often recommend different vendors for the same question on the same day, so each platform needs its own tracking.

The study

We reused the 40 B2B software buying questions from our Claude tracking study, two each in 20 categories such as CRM, payroll and endpoint security. From October 3 to October 9, we asked every question once a day on ChatGPT, Gemini and Claude (Opus 5.5 through the API). To check that the results hold for people using claude.ai itself, we added answers a person had collected by hand in claude.ai for the earlier study. In all, 2,770 answers were evaluated in this study, including a side test on older ChatGPT answers to consumer and agency questions.

For every answer, a classifier labeled each vendor it named: the top choice, one of several good options, a mention with no recommendation, or a vendor the answer cautioned against. It's the same classifier Spyglasses uses in production. We call the recommended vendors, top choice or one of many, the shortlist.

Then we asked a simple question. When a vendor is on the shortlist in one run, is it on the shortlist again in another run of the same question? We set the tests and their margins before collecting any data.

What we found

The shortlist repeats

Dot plot with one row per platform. The share of shortlisted vendors that were shortlisted again in another run of the same question is 81% for ChatGPT, 78% for the Claude API, 75% for claude.ai and 64% for Gemini, each with a 90% interval.

When ChatGPT shortlisted a vendor, that vendor made the shortlist again 81% of the time in another run of the same question. The Claude API came in at 78%. claude.ai, asked by hand, was close at 75%. Gemini trailed at 64%.

Most of each shortlist is a stable core. On ChatGPT, 60% of all the shortlist places across a question's seven runs went to vendors that were shortlisted in every single run. On Claude it was 56%. Vendors that made the list only once were rare, 4% of places on both. Gemini's core was smaller, at 32%, with more vendors coming and going.

Stacked bars for ChatGPT, the Claude API and Gemini, splitting every shortlist place by how many of the question's seven runs that vendor was shortlisted in. In every run, 60%, 56% and 32%. In five or six runs, 19%, 19% and 31%. In two to four runs, 17%, 21% and 28%. In one run only, 4%, 4% and 9%.

Part of that stability comes from the market itself. Every category has well-known names, so two different questions in the same category share some of their shortlist (59% on ChatGPT). The same question repeats much more (81%). The question your buyer asks decides the rest.

The kind of market matters too. In the side test on older ChatGPT data, consumer headphone questions repeated their shortlist about as often as the B2B software questions. Brand-design agencies, a crowded market of many small firms, repeated theirs far less, though that sample was small.

One more check: the shortlist turned out to be exactly as stable as the plain list of every vendor an answer names. It just tells you more, because it separates the vendors AI endorses from the ones it only mentions.

"Best for" moves more

Answers often say what a vendor is best for, such as small teams, enterprises or a particular industry. That framing was less stable than the shortlist. A vendor paired with the same use case came back about half the time on ChatGPT and Claude, and less often on Gemini. Why the framing moves while the shortlist holds is a question we'd like to study next.

The platforms disagree with each other

Dot plot. Each platform asked again on another day overlaps with its own shortlist at 81% for ChatGPT, 78% for the Claude API and 64% for Gemini. Two platforms on the same day overlap at 63% for ChatGPT and Gemini, 63% for ChatGPT and the Claude API, and 55% for Gemini and the Claude API.

For the same question on the same day, ChatGPT's and Gemini's shortlists overlapped 63%. ChatGPT and Claude also overlapped 63%, and Gemini and Claude 55%. ChatGPT and Claude each agreed with themselves on another day more than with any other platform on the same day. Gemini agreed with itself about as much as it agreed with ChatGPT. Your place on one platform's shortlist doesn't tell you where you stand on the others.

A short note on rank and the top pick

A vendor's exact position repeated far less often than its place on the shortlist: 35% of the time on ChatGPT, 27% on Gemini and 22% on the Claude API. The top pick repeated 55% of the time on ChatGPT and less often elsewhere. But a top pick that lost the top spot rarely left the list. On ChatGPT and Claude, the top pick from one run was still shortlisted in the other run 95% to 99% of the time. Some of that rotation comes from the labeling itself. When we re-labeled identical answers, the top pick changed more often than the shortlist did.

This matches our earlier research on rank. A brand slipping from first to third on one question is normal day-to-day movement. A brand leaving the shortlist is a change worth acting on.

What this means for your tracking

Make the shortlist your goal. Being recommended for your buyers' questions is a stable, trackable result. A specific position isn't. Set goals and report wins on shortlist inclusion.

Don't report a rank change from one run. A move from first to third is well within normal day-to-day movement. Look at how often you're recommended across many runs before you call it a trend.

Track each platform on its own. ChatGPT, Gemini and Claude disagree often enough that a combined number hides real differences. Gemini needs more runs before you trust a change.

Read "best for" across runs. If you care which use case AI ties you to, look at the pattern over time, not a single answer.

What this doesn't mean

One question isn't a category, and these were 40 synthetic B2B software questions asked from a US location over seven days in a row. ChatGPT and Gemini were collected through a third-party data service, Claude through the API, and claude.ai by hand from one account. The labels come from an AI classifier. The study says nothing about what buyers do with the answers. ChatGPT also switched to a newer model on the sixth day, which leaves fewer days per model. Looking at the first model alone, its shortlist repeated 83% of the time, and the conclusions don't change.

What's next

Seven days is a short window. Stage 2 asks the same 40 questions on the same three platforms once a week through November 6, at the same time of day. It tests whether the shortlist still holds when the runs are about a month apart instead of a day or two. The plan and the tests are already set, and we'll publish the results in November.

Track the measure that holds

That shortlist measure is what Spyglasses tracks. Daily Prompt Tracking checks every day whether each platform recommends you, as the top choice or one of many, on the questions your buyers ask. And because the platforms disagree, it reports each one separately, on ChatGPT, Gemini and Google AI Overviews, with Claude available as an add-on.

Getting Shortlisted by AI Is Stable. Your Rank Isn't.