We Matched Synthetic Prompts to Real Human Intent. The Brand List Transferred. The Percentages Didn't.

Jim Wrubel
8/6/2026

Four days ago we published a study on our own prompt generator and ended it with a bet. If a synthetic prompt panel were built to match how real buyers actually phrase things, sub-intent by sub-intent, the data said it should mirror a human panel at every level. We pre-registered that test, froze the panels before collecting anything, and committed to publishing whatever came back.
Half the bet paid. The half that failed turned out to be the more useful result.
Key takeaways
- Synthetic prompts matched to real human sub-intent produced answers that are statistically indistinguishable from human prompts, overlapping at 0.516 against a human to human baseline of 0.517.
- That same panel still missed the human brand share vector by 8.9 percentage points, against the 5 point equivalence band registered in advance.
- The disagreement is about frequency, not membership. Both panels surfaced the same six brands in nearly the same order, with a rank correlation of 0.87.
- Within a single use case, 82 percent of human prompts and 72 percent of synthetic prompts surfaced Apple at least once, but 9.5 percent versus 0 percent surfaced it in all five daily runs.
- Letting a brand name its own use cases does not repair the percentage. Restricting the comparison to prompts sharing a use case left the gap at 0.085, against 0.089 unrestricted.
- Synthetic prompts are as self consistent as human prompts when re run, so the gap is systematic rather than random noise.
- Accuracy depends on rank. The two best known brands came back within 6 percent of their real numbers, the fifth and sixth were off by 33 and 46 percent, and several smaller brands never appeared in the synthetic panel at all.
What we did
We took 143 prompts that real people wrote for one buying intent, headphones as a travel gift, from SparkToro's survey data. Then we generated a synthetic panel built to match that human panel's profile across six dimensions of how people ask: travel context, music use, budget limits, named recipients, form factor, and wireless. A second synthetic panel got no matching at all, as a control. All three ran against ChatGPT every day for five days. 1,230 runs evaluated in this study.
The hypotheses, the equivalence bands, the panels, and the analysis code were frozen before we collected a single answer. The full study is on our research site, including the complete text of both synthetic panels so anyone can re run them and check us.
What we found
The prompts themselves are now interchangeable with human ones. Pick a synthetic prompt and a human prompt at random, compare the two answers, and they agree as much as two differently worded human prompts agree with each other. The overlap was 0.516 against a human baseline of 0.517. Our equivalence band was 0.10 and the difference came in at 0.001. This is the closest match to human phrasing we have measured, and it beat our own prediction.
The panel's percentages still missed. Averaged across the six brands humans mention most, the matched panel sat 8.9 points away from the human panel. Better than the unmatched control at 13.7, better than last month's neutral panel at 11.2, and still outside the 5 point band we committed to. The prediction failed and we are reporting it as a failure.
But "the percentages missed" describes something narrower than it sounds. A share figure pools a yes or no per answer, so being off by 8.9 points can mean two very different things. Either the panel surfaces different brands, or it surfaces the same brands at different rates. We split the statistic to find out.
Within the travel use case, the panels agree on which brands belong. Counting prompts that surfaced each brand at least once across their five runs: Sony 99 percent of human prompts and 96 percent of synthetic, Bose 95 and 96, Sennheiser 94 and 89, Anker 91 and 100. Even the worst case, Apple, was 82 against 72. Both panels returned the same six brand consideration set in nearly the same order.
Where they part company is how often. Apple came back in all five runs for 9.5 percent of human prompts and none of the synthetic ones. Sennheiser, 59 percent against 22. The two panels see the same market and report it at different volumes.
How far off the percentage runs depends on where you sit in the ranking. The panel is nearly exact at the top and drifts badly toward the bottom. Measured against each brand's own human number:
| Brand | Rank in the human panel | How far the synthetic panel was off |
|---|---|---|
| Bose | 2 | 0.6 percent |
| Sony | 1 | 6.4 percent |
| Anker | 4 | 12.0 percent |
| Sennheiser | 3 | 13.2 percent |
| JBL | 6 | 33.1 percent |
| Apple | 5 | 45.8 percent |
Below those six it gets worse. Several brands that real prompts surfaced in 2 to 4 percent of answers never showed up in the synthetic panel at all, including Bowers and Wilkins, Sonos, Focal, and Beyerdynamic. Two went the other way, with Shokz appearing 2.6 times more often in the synthetic panel than real prompts produced.
That direction matters for how you read your own number. It is not that smaller brands get undercounted. It is that below the top few, the number stops being reliable in either direction.
Naming your own use cases will not fix the percentage. The intuitive workaround is to have a brand list the use cases it cares about and generate prompts for those, sidestepping the need to match a mix at all. We tested it by comparing only prompts that carried the same use case. The gap did not move: 0.085 within travel, against 0.089 for the unrestricted comparison. You get a realistic brand list either way, and you do not get a trustworthy percentage.
Synthetic prompts are not flaky, which cuts against what we published last month. Re run five times, a synthetic prompt is as self consistent as a human one. Our previous study measured all three of its synthetic panels as materially noisier than the human panel and concluded that synthetic prompts occupy a less stable region of the response space. This study could not replicate that. The two studies' intervals overlap, so this is a failure to replicate rather than a refutation, but the general claim does not survive. We are reporting it because a research program built on registering predictions in advance has to publish the numbers that cut against it first.
What this means for reading your dashboard
- Read the ordering, not the percentage. If a share of voice figure comes from a synthetic panel, the brand list and the rough ranking are the parts this evidence supports. The percentage is softer than it looks.
- The consideration set is the reliable output. "Who am I competing with for this use case, and roughly where do I stand" is a question synthetic panels answer well, and it is enough to drive content planning and competitive positioning.
- Ask what your panel never asks. Our matched panel hit all six dimensions it was told to match and stayed silent on nine others. It never asked about watching films on a plane, which appeared in 31 percent of human prompts, or the recipient's age, or how many options to return. It reproduced 13 percent of the distinct ways real people combined those dimensions. Coverage of the combinations your buyers actually use is the thing worth interrogating.
- Be most careful if you are not one of the category leaders. The two best known brands in our study came back within 6 percent of their real numbers. The fifth and sixth were off by 33 and 46 percent, and brands below them were often missing entirely or, in two cases, substantially overstated. If you are the challenger in a category, your specific number is the least reliable one on the page, and it can err in your favour as easily as against you.
What this changes in Spyglasses
We ran this study on our own methodology, so the obligations land on us first.
- We are separating the two questions in the product. "Which brands compete for this use case, and where do you rank" is what the evidence supports, so that is what gets the prominent treatment. Percentage share of voice keeps its place, with the ordering carrying the interpretation.
- Phrasing coverage becomes a reported number, not an internal one. The share of distinct buyer phrasing profiles a panel reproduces was 13 percent here. That is a measurable property of any panel and we would rather show it than have it sit unexamined.
- We publish our own instrument's errors. The audit for this study found our brand extractor counts brands used as platform references, which inflated the single largest gap in the article by roughly 18 percent, in the direction that made our own headline look better. It is in the study, with the corrected numbers.
Read the full study, the pre-registered predictions, and both datasets.