We Tested Our Own Synthetic Prompts Against 143 Real Humans

Jim Wrubel
8/2/2026

Every AI visibility tool, ours included, tracks brands by running prompts that a person never typed. Critics have a fair question about that: if no human wrote these prompts, what exactly do the numbers describe? A widely shared list of 31 methodology questions for vendors put the burden of proof on us. So did SparkToro's research on how wildly real people vary their phrasing.
We think the burden of proof argument is right. So we pointed the test at our own product.
Key takeaways
- A prompt panel generated from a brand's own website measures that brand's home field, the buying situations its positioning claims. It is not a measurement of the overall market.
- The most reliable signal from a brand-anchored panel is losing on it. Winning there is the expected outcome; a competitor beating you on prompts built from your own positioning is the finding to act on.
- Competitive gaps read off an anchored panel are biased against your rivals, not inflated for you. Swapping the anchor brand moved the measured gap between two competitors by 41 points.
- Scenario-generated synthetic prompts produced answers statistically indistinguishable from real human prompts, and when a synthetic and a human prompt shared the same sub-intent, their answers matched completely.
- Matching a synthetic panel's sub-intent mix to real buyer behavior is a testable recipe for market-representative tracking. We pre-registered that test and it is collecting data now.
What we did
We took 143 real prompts that real people wrote for one buying intent (headphones as a travel gift, from SparkToro's survey data) and ran them against ChatGPT every day for five days. Alongside them, we ran three synthetic panels: two generated by our own production prompt generator pointed at bose.com and soundcore.com, and one generated from the buying scenario alone with no brand involved. Same platform, same days, 1,325 runs evaluated in this study.
The hypotheses, the statistical bands, and the analysis code were all frozen before we collected a single response. Whatever came out, we committed to publishing it. The full study is now live on our research site, including the complete text of every synthetic prompt so anyone can re-run the panels and check us.
What we found
A prompt panel generated from your website doesn't measure the market. It measures your home field. The generator builds prompts from what your site promotes: your features, your segments, your product lines. Those prompts ask the questions your positioning claims to win. When we reweighted the human panel to ask the same mix of questions the Bose panel asked, the humans saw nearly the same market the panel saw. The panel isn't broken. It's answering a narrower question than the label on the dashboard suggests.
The home field doesn't flatter you, and that's the useful part. Neither anchored panel inflated its own brand above what real humans saw. And Bose still lost to Sony on every panel we ran, including the one generated from bose.com. That asymmetry is the practical takeaway: winning on your own panel is the expected outcome. Losing on your own panel is the alarm that should reorganize your content plan.
Competitive gaps from an anchored panel are the number to distrust. Swapping the generator's anchor from one brand to its rival moved the measured gap between them by 41 points. Not because either panel puffed up its own brand; because each panel failed to surface the rival. If your dashboard says you're beating a competitor on prompts anchored to your own site, the competitor's number is the one most likely to be wrong.
And the good news: there's a real path to synthetic panels that mirror humans. The neutral, scenario-generated panel produced answers statistically indistinguishable from human-prompt answers. Its one remaining failure was the mix: it over-sampled some sub-intents and missed others, like budget limits and named recipients. When we compared only prompts that shared a sub-intent, synthetic and human answers matched completely. Authorship didn't matter. The sub-intent did.
That last finding is a recipe, not just a result. If you generate scenario prompts and match their sub-intent mix to how real buyers actually ask, the data says the panel should mirror the human one at every level. We've already launched the follow-up experiment to test exactly that, with the predictions registered publicly before any data arrives.
What this changes in Spyglasses
We ran this study on our own generator, so the obligations land on us first:
- Share of voice from brand-anchored prompts gets labeled as what it is: your standing in the conversations your positioning claims. We already exclude awareness-stage prompts from share of voice; the label makes the rest of the story explicit.
- Panel generation is moving scenario-first. The neutral generation frame is the one that survived contact with real humans, so it becomes the default path, with brand anchoring reserved for the home-field question it actually answers.
- Phrasing diversity beats prompt volume. Real people state budgets, name the recipient, and ask for "top 5 with reviews." Panels that never do those things sample a narrower market no matter how many prompts they run. We're building those behaviors into generation rather than adding more prompts that ask the same question.
This is also why we've always said measurement is the proof layer, not the work itself. Prompt tracking tells you where you stand on the questions you chose to ask. Moving the answer happens in the retrieval pipeline: the searches AI runs, the sources it can read, the pages it cites. That part is measurable and improvable regardless of which panel you track.
Check our work
The full study has every number, the pre-registered spec, both datasets, and the complete synthetic prompt text under a CC BY license. If you work at another AI visibility vendor: the human baseline is obtainable, the method is documented, and we'd like to see your number next to ours.