How People Word a Prompt Changes Whether AI Recommends You

Jim Wrubel

Jim Wrubel

7/23/2026

#prompt tracking#AI visibility#ChatGPT#brand mentions#research#AEO
How People Word a Prompt Changes Whether AI Recommends You

Summary: Spyglasses Research extended SparkToro's January 2026 prompt-consistency study by re-running their 143 de-identified, human-written headphone prompts against ChatGPT (via DataForSEO's LLM scraper, en-US, web search on) once a day for seven days. Category leaders Sony, Bose, and Sennheiser appeared in 80 to 90% of answers regardless of wording, replicating SparkToro's finding on a single platform. But prompt phrasing changed the rest of the recommendation set more than ChatGPT's own run-to-run randomness: same-prompt reruns overlapped at 74% (Jaccard 0.74) while differently worded same-intent prompts overlapped at 54%. Much of that gap traces to sub-intents. A stated dollar budget dropped Bose 36 points and lifted JBL 31; travel framing lifted Anker 23; music framing lifted Sennheiser 23. For measurement, phrasing variety beat run frequency: 70 varied prompts run once estimated true mention rate about twice as accurately as 10 prompts run daily for a week.

Key takeaways

  • Market leaders are close to wording-proof; mid-tier brands live or die by how the question is phrased.
  • Rewording a same-intent prompt swaps roughly one brand per shortlist, beyond the model's own noise.
  • Concrete prompt attributes (budget, recipient, use case) act like separate markets with separate leaderboards.
  • A fixed prompt panel measures drift well but can't tell you your true mention rate; a periodic broad sweep can, and it audits the panel too.

Ask ChatGPT for headphone recommendations the way you'd word it, and the way your neighbor would word it, and you'll often get a different shortlist. Not a different market leader; a different shortlist. That gap is the single most useful thing we learned from our latest study, and it changes how you should think about every AI visibility number you look at.

Quick background. In January 2026, SparkToro published a study built on a great idea: instead of guessing what "the typical prompt" looks like, they asked over 142 real people how each of them would ask an AI for headphone advice, ran every prompt, and compared the answers. Same brands kept showing up. But each prompt ran once, spread across seven different AI platforms, so wording effects and platform effects were tangled together.

With their blessing (thanks, Rand), we took their exact prompts and untangled it: one platform, ChatGPT, every prompt re-run daily for a week. 1,001 headphone answers evaluated in this study, plus a control set. Full methods and data are on our research site.

What held steady, and what didn't

The leaders are close to wording-proof. Sony appeared in 90% of answers, Bose in 82%, Sennheiser in 80%, across 143 different people's phrasings, every day, all week. If you're the category leader, the "is my prompt representative?" worry mostly isn't your worry.

Below the podium, it's a different story. We went in expecting wording not to matter once intent was the same; we even pre-registered that prediction. The data said otherwise. Two runs of the same prompt agree on about 74% of recommended brands. Two differently worded prompts, same intent, same day: 54%. Wording costs about one brand per shortlist, and that's after accounting for ChatGPT's natural run-to-run variation.

Here's the part you can act on. Much of that "wording effect" isn't mysterious at all. It's sub-intent: concrete attributes people bake into how they ask.

Grouped bar chart comparing brand mention rates for prompts with and without a specific dollar budget. Bose drops from 88% to 52% and Sennheiser from 83% to 66% when a budget is stated, while JBL rises from 12% to 43%.
Prompts that name a specific dollar budget get a different market. Bose drops 36 points, JBL more than triples.

When someone names a dollar amount, Bose drops 36 points and JBL gains 31. Mention travel, and Anker climbs 23 points. Frame it around music listening, and Sennheiser climbs 23. These aren't vibes; each of those effects held up under the same statistical checks as the main study. Some didn't, and we say so in the research article: mentioning noise cancelling moved nothing (it's table stakes in this category), and asking for a specific output format was too rare in the prompt set to test.

So "prompt wording matters" decomposes into something a marketer can actually use. A budget-framed question is a different market than a premium-framed one, with a different leaderboard. You can't enumerate every way a customer might phrase a question. You can absolutely enumerate the frames: price-conscious, gift-buying, use-case-specific, feature-led. And you can find out which of those frames include you and which shut you out.

And these segments aren't just a pattern in the answers. We checked the layer underneath, and prompts that share a frame also run more similar web searches inside ChatGPT and cite more similar sources. Two travel-framed prompts overlap far more in their grounding searches than a travel prompt paired with a non-travel one. For budgets, the segment turns out to be the amount itself: prompts naming similar dollar figures converge on similar brands, while a $100 prompt and a $400 prompt behave like two different markets. The segments are real all the way down the pipeline, which is what makes them worth targeting.

What this means for your tracking

If you track AI visibility with a fixed set of prompts (most tools work this way, ours included), your mention rate is a precise measurement of those exact phrasings. Run them more often and the number gets more stable. It doesn't get more representative, because the wording effects are stable too. A panel that happens to over-sample budget-framed wording will understate a premium brand every single day, consistently, convincingly.

Line chart showing the margin of error on estimated mention rate versus the number of distinct phrasings in a panel. Error falls steeply as phrasings are added; running the same phrasings daily for a week improves accuracy only modestly despite seven times the cost.
Simulated from our study grid: adding differently worded prompts improves accuracy fast. Re-running the same prompts all week barely does.

In our simulation, 70 differently worded prompts run once pinned down a brand's true mention rate about twice as tightly as 10 prompts run daily for a week. Same total cost. Variety wins.

Does that mean scheduled tracking is a waste? No, and this matters: our study also found zero day-to-day drift across the week. A fixed panel is a clean yardstick for change. When a model update lands or your campaign ships, the panel is what tells you something moved. It's the right tool for exactly that job, and the wrong tool for estimating your overall presence.

The design that follows, and the one we're building toward, is what we call a pulse model. Keep a small, fixed panel running on a schedule to catch drift. Periodically, run a broad sweep: a large, one-shot batch of prompts deliberately varied across the frames above. The sweep tells you your true rate per segment and which frames exclude you (that's your content and PR target list). It also calibrates the panel: if the sweep says 45% and your panel says 60%, you now know your panel reads 15 points hot, and every daily number after that gets more honest.

Limitations of our study

This was one buying intent (headphones as a travel gift), one platform, in English, over one week. It also covers exactly one kind of user: an anonymous, free-tier session with no personalization, which is what the scraper sees. Logged-in accounts, paid tiers, and personalized histories could all behave differently, and we can't claim these numbers carry over to them. The pattern of leaders-stable-and-shortlists-moving is the kind of mechanism we'd expect to travel to other categories, but we haven't proven it travels. The sub-intent effects and the panel simulation are exploratory analyses, clearly labeled as such in the full write-up. Don't stretch them further than we do.

The full methodology, the pre-registered spec, every statistical test, and a downloadable dataset live on our research site: We pre-registered "prompt phrasing barely matters." ChatGPT proved us wrong.

Find out which frames include you

The first question this study should make you ask isn't "what's my mention rate?" It's "which ways of asking include me, and which don't?" The Spyglasses AI Visibility Report is the fast way to start: it runs your brand across ChatGPT, Gemini, Claude, and Perplexity and shows you what they say, what they cite, and how consistent it is. From there, you'll know whether the wording gap is costing you answers you should be in.