A Modest Defense of Prompt Tracking

Jim Wrubel
7/29/2026
There's a growing genre of articles from respected names in SEO and marketing that are skeptical of AI visibility tools. Some are measured. Some are outright skeptical. One is even from us. The most useful one so far is a recent piece posing 31 methodology questions for prompt tracking platforms, on the theory that vendors are making the positive claim, so vendors carry the burden of proof.
That's correct. We do.
So this is a defense of prompt tracking, but a modest one. Modest because the critics are right about a lot. There are snake oil sellers in this category. There is no industry standard for how data gets collected or reported. There are dashboards showing three-point share-of-voice movements as wins to clients who have no way to know the number's margin of error is wider than the movement. If you sell AI visibility software and none of that stings, you're not paying attention.
But "this industry has real problems" and "this measurement is worthless" are different claims. The first is true. The second doesn't jive with how every other measurement tool marketers trust actually works.
Key Takeaways
- The data sources SEO professionals trust (SERPs, Search Console, GA4, Ahrefs, Semrush) can be served cheaply at scale because they're cached. AI answers can't be cached, so every observation costs a fresh generation. That's a structural difference, not a maturity gap, and it's a big part of why these tools are expensive, which is another frequent criticism.
- Those trusted sources are also built on sampling and modeling. The difference is they've had years to normalize their caveats.
- Nobody in this industry has first-party prompt data, and nobody should want the versions of it that are for sale.
- We scored ourselves against all 31 methodology questions: some answered with published research, some being tested right now, some answered by disclosure, and some that can never be claimed by anyone.
Cache rules everything around me
Start with the question underneath the whole debate: why does AI visibility data feel so much shakier than the SEO data everyone already trusts?
Because the SEO data you trust is cached.
A SERP is generated once and served millions of times. Google Search Console hands you a pre-aggregated export of things that already happened. GA4 processes events into tables long before you open a report. Ahrefs and Semrush crawl, process, and store the web's link graph and ranking data so that your query hits a database, not the live web. Caching is what makes those products cheap, fast, and consistent. Ask the same question twice, hit the same stored answer twice, and the data feels solid.
AI answers don't work that way. Every response is generated fresh, per request, at real compute cost. There is no stored answer to check your brand against. The only way to observe how an assistant talks about your category is to ask it, pay for the generation, and record what came back. Ask again and you may get something different, because the system itself is nondeterministic.
This has two consequences the critics usually treat as vendor failures when they're actually physics:
- Every AI visibility number is a sample. There's no census to compare against, because the population of "answers ChatGPT would give" doesn't exist until each answer is generated.
- Scale costs real money. A rank tracker can check 10,000 keywords for pennies against cached indexes. Ten thousand AI prompt runs are ten thousand paid generations, per platform, per day.
So AI visibility tools sample. The defensible ones say so, publish the math, and show you the confidence interval. The indefensible ones report a single daily run as if it were a fact. The difference between those two isn't the data source. It's the discipline.
The tools you trust are sampled too
The cached sources our industry relies on are also samples, and often heavily modeled ones.
GA4 applies sampling to complex reports and thresholding to small segments; it also models conversions and behavior where consent or identity gaps exist. Search Console filters low-volume queries for privacy and caps exported rows; the "total" you see is not the total that happened. Ahrefs and Semrush estimate search volume and traffic from clickstream panels and models; two tools will confidently give you two different "monthly search volumes" for the same keyword, and neither is an observation. It's an estimate built on a panel that skews toward whoever installed the panel's software.
None of this makes those tools bad. They're useful because they're directionally right, methodologically stable, and increasingly upfront about their limits. But the industry's trust in them was earned over years of practitioners learning where the numbers bend. Nobody writes "31 methodology questions for rank trackers" anymore, not because rank trackers answered them all, but because the caveats got absorbed into professional common knowledge.
AI visibility measurement is at the beginning of that same curve, with a harder problem (no cache, nondeterministic output) and less accumulated trust. The right response to that isn't to declare the category unmeasurable. It's to move through the curve faster by publishing methodology and running the studies critics keep asking for.
The first-party data nobody has
The strongest single objection in the skeptical articles is this one: no vendor has first-party data on what real users actually type into AI assistants, what answers they receive, or what the assistant's memory knows about them. Panels of tracked prompts are a stand-in for a population nobody can observe.
We're not going to argue with that. It's true, it will stay true for as long as this industry exists, and any vendor implying otherwise should be asked where their data comes from.
That last question matters more than it seems, because there is prompt data for sale. It comes from free VPNs, browser extensions, and "free" analytics tools that record what their users type into AI assistants and resell it. Some tools in our category buy it and present the result as observed demand.
We won't. Not because the data would be useless, but because of how it's gathered. People typing into a chat box have a reasonable expectation that the conversation stays between them and the assistant. Harvesting those conversations through a free utility whose users never meaningfully agreed to be a data source is not a data strategy we'll participate in, as a buyer or a seller. Our data provenance commitment is written into our methodology docs: every prompt we run is one you explicitly asked us to run, the raw answers are inspectable and exportable, and we don't buy harvested conversations or inflate seed prompts into fake "demand" data.
If that means our picture of user behavior is a panel rather than a population, we'll take the panel. At least we can tell you exactly what's in it.
Scoring ourselves against the 31 questions
The 31 questions deserve better than a vibes-based response, so we mapped every one of them against where we actually stand. Four buckets came out of that exercise.
Answered with published research
Some of the questions are empirical, and we've already run the studies. Our research site publishes them with pre-registered specs, frozen before data collection, with the raw derived datasets released under CC BY.
- "What's the run-to-run variance? When is one run per prompt enough?" (Q8, Q18): We re-ran 143 human-written prompts against ChatGPT daily for a week and measured exactly how much the same prompt's answer moves between runs, at every layer: brands recommended, sources cited, searches run.
- "Does daily movement mean anything or is it noise?" (Q12): Same study. Across seven daily waves we found no day-to-day drift in the underlying retrieval layer; the variance is run-to-run, not date-to-date.
- "Doesn't phrasing determine the result?" (parts of Q3 and Q7): Yes, and we published that finding against our own expectation. We pre-registered the hypothesis that phrasing wouldn't matter beyond noise. The data rejected it. Rewording a same-intent prompt moves brand recommendations more than the model's own randomness does, and we said so, because the point of pre-registration is that you don't get to bury the result you didn't want.
Being tested right now
The hardest questions in the article are about representativeness: do synthetic prompt panels resemble what real people ask (Q3), does panel configuration predetermine the result (Q6), does more volume add validity or just precision (Q7), and what does generated prompt text actually represent (Q17)?
As this post goes out, we have a pre-registered experiment in the field testing exactly those four questions, using our own prompt generator as the test subject, against the hardest baseline available: real human phrasings of the same intent, collected the same week, on the same platform. The spec was frozen in a public repo before the first response came back, which means we've already committed to publishing the result whether it flatters our product or not. We've done that before (see the phrasing study above), and it's the only version of "trust us" we think this industry should accept.
Cross-vendor agreement (Q19) is on the same roadmap, and it really should be an industry exercise. If any other vendor wants to run the same prompts, same week, and compare outputs in public, our inbox is open.
Answered by disclosure
A large cluster of the questions aren't research questions at all. They're "show your work" questions: what exactly counts as a mention, how is share of voice calculated, what's the unit of counting, what counts as a citation, what happens when a model refuses, can customers audit the raw data, what happens to history when methodology changes (Q1, Q5, Q20 through Q26).
These get answered by documentation or they don't get answered at all, so we publish the entire collection and calculation methodology: which platforms run through anonymous logged-out sessions versus APIs and why cross-platform comparison is therefore asymmetric, the exact formula behind every headline metric, why our share of voice is a presence rate that deliberately doesn't sum to 100%, how mention detection stays deterministic so any number we show can be found in the raw answer text, how alias changes recalculate your full history instead of leaving a silent discontinuity, and which signals we refuse to ship because we can't make them stable (sentiment on regenerated answers, cross-platform position averages).
One disclosure worth calling out, because it answers the article's uncertainty questions (Q9, Q10) directly: we show you the margin of error before you spend anything. When you pick prompts to track, the platform displays the observation count, the confidence interval, and a plain quality rating while you're still choosing. Ten prompts tracked daily gives you roughly ±4.7% at 90% confidence; three prompts gives you ±8.7%, and we label that directional rather than letting you present it as precise. The calculator and formula are public.
What nobody can claim
And then there's the bucket where the only defensible answer is "no one can know this, including us."
- Personalized, logged-in behavior (Q16). We measure logged-out, memory-free sessions because they're the only measurable and comparable baseline. Real users have history, memory, and personalization, and no vendor can observe those sessions. Logged-out is a proxy, not a mirror, and our docs say so in those words.
- Population-level representativeness (Q2). Without first-party usage data, no panel can be proven to represent "the market." A tracked panel measures the panel. Movements within it are evidence about the panel first and the world second.
- Causal chains to revenue (Q27, Q29). "Your visibility score went up, therefore the content change worked, therefore revenue" is a chain nobody in this category has demonstrated with controls. We do have one advantage here, as do several other, larger platforms in our space: our traffic analytics detect actual AI assistant visits server-side, which gives brands a first-party outcome to correlate against instead of a score validating itself. But correlation with controls is a study we still owe, not a claim we get to make today.
The modest defense
Prompt tracking is a legitimate instrument with a narrower job than most of its sellers admit and more real utility than its critics allow.
Run a fixed panel on a schedule and you have a drift detector: it tells you when something materially changed. Might be a model update, a competitor move, or your own work changed how assistants answer in your category. Even if the source isn't something we can infer and report, it's a signal to investigate.
Pair it with periodic broad sweeps of varied phrasings and you can measure your actual presence and audit whether your panel runs hot or cold. Ground it in the retrieval layer (the searches assistants actually run, and where you rank in them) and the noisy output metric gains a deterministic layer underneath that you can act on.
What prompt tracking is not: a census of user behavior, a mirror of personalized sessions, or a benchmark of market truth. Any vendor presenting it that way is exaggerating at minimum.
The critics asking hard methodology questions are doing this industry a favor. The snake oil they're reacting to is real, and the fastest way to separate from it is boring: publish your methodology, pre-register your studies, release your data, state your margins of error before the invoice, and put the limitations in writing where clients will actually read them.
That's the standard we're holding ourselves to. We'd be glad if it became the industry's.