How to Track Your Brand's AI Visibility Over Time

Jim Wrubel
7/30/2026

A new client signs, you run the baseline report, and the deck looks great. Then three weeks pass and someone asks: "is it working?" A single report can't answer that question. It's a snapshot, not a trend line, and a trend line is the only thing that proves visibility moved because of the work and not because AI gave a different answer on a different day.
That's the gap between a report and tracking. A report tells you where a brand stands right now. Tracking tells you whether it's getting better, and whether you can point at what caused it. This guide walks through setting that up, step by step. It applies to any AI visibility platform; the callouts show how each step works if you run it in Spyglasses.
| Step | What you're doing | What tells you it's working |
|---|---|---|
| 1. Confirm your baseline | Fixing the starting point | One completed report, reviewed for accuracy |
| 2. Revise competitors | Measuring against the right list | 4+ competitors the client actually names |
| 3. Expand prompt coverage | Covering the buyer journey | Coverage score above 60 |
| 4. Decide the messages | Defining what "on-message" means | Positive and negative claims captured |
| 5. Turn it into a project | Getting a nightly cadence | Project active, margin of error reliable |
| 6. Schedule reports | Automating the recap | Recurring schedule with recipients set |
| 7. Watch the trend | Reading movement correctly | Slope tracked over weeks, events annotated |
Confirm your baseline is solid
Everything downstream compares back to the first report, so it's worth a second look before you build on it. Skim the summary and the per-platform breakdown across ChatGPT, Gemini, Claude, Perplexity, and Google AI Overviews, plus the recommendations. If the report crawled and generated a brand snapshot automatically, check that the category, positioning, and competitor list actually match how the client describes themselves. A wrong category quietly skews the prompts and mention matching that come next.
This is also the moment to decide what "improved" will mean for this client. Is the goal more mentions, a better share of voice against a named rival, or cleaner brand consistency across platforms? Write it down. You'll want it in three weeks when someone asks whether the numbers moved.
Revise your competitor set
Auto-seeded competitor lists come from whoever gets co-mentioned in the first batch of answers. That catches some real rivals and some noise; a review site that happened to come up once, a brand from an adjacent category, a company the client hasn't thought about in years. None of that is who you're actually measuring against.
Go through the list with the client's own reporting in hand. Prune what doesn't matter, add the names they use in their board deck, and set aliases for companies that go by more than one name or domain. Share of voice is only useful if it's measured against the competitors the client cares about.
Expand your prompt coverage
Five or six starter prompts are enough to prove a first report works. They're not enough to track anything. Real coverage spans the questions a buyer actually asks across a full journey: brand-identity questions ("who is Acme Outdoor"), comparison questions ("Acme vs. SummitGear"), category questions ("best tents for alpine conditions"), and decision-stage questions ("should I buy from Acme").
Build the set with a mix of methods. Generate prompts from a framework built for this (category entry points, jobs-to-be-done, buyer's-journey stages), import any prompt list the client already tracks elsewhere, and fill obvious gaps by hand. Aim for coverage across all four question types, not just volume in one.
Decide what messages you want to see win
Share of voice tells you how often a brand gets mentioned. It doesn't tell you whether AI is saying the right things about it. That's a separate question, and it needs its own input: the specific claims the client wants to see repeated, written down as short, concrete statements instead of a general sense of "good coverage."
Positive key messages might be a differentiator, a guarantee, or a category claim the client wants to own. Negative ones are anything the client wants to catch early, an outdated claim, a discontinued product, an old executive's name. Once these are captured, matching is semantic; a paraphrased version of the message still counts as a hit, so you don't need to predict the exact wording AI will use.
Turn tracking into a running project
None of the above runs on its own. Prompts only get tested on a schedule inside a project, so this is the step that actually turns a static list into tracking. Pick a project type that matches the goal: an SEO project if you're measuring the effect of new content or site changes, a PR project if you're measuring earned media, or a general project if you just want the trend line.
Auto-select prompts from the set you built, then check the margin of error before you move on. A project running 12 prompts nightly across ChatGPT, Gemini, Claude, Perplexity, and Google AI Overviews produces a much tighter confidence band, often under 8%, than five prompts run once, which can sit closer to 15%. That gap is what separates a real change from noise in next month's report.

“A report tells a client where they stand today. A project run nightly is what tells them whether the work is moving the number.” — Spyglasses
Put reports on a schedule
Tracking that only you can see doesn't help a client relationship. Automate the recap so it lands in an inbox on its own, weekly, monthly, or quarterly depending on how the engagement is structured. Add the client's actual stakeholders as recipients, not just the day-to-day contact; visibility data has a way of becoming a board-level topic once it's been running a few months.
This also protects the relationship from the "what have you been doing" conversation. A recurring report in someone's inbox is proof of ongoing work even in a month where nothing dramatic happened.
Watch the trend, not the daily noise
Once nightly data starts flowing, resist the urge to react to any single day. AI answers are nondeterministic; the same prompt can come back with a different answer on the same platform two nights in a row. One good or bad night is not a signal. A slope held over two or three weeks is.
Two habits make the trend readable. First, annotate anything that might move the numbers: a new page published, a press hit, a competitor's rebrand. Second, check the trend against the margin of error before calling a change real. A three-point jump in share of voice matters a lot if your confidence band is one point wide, and means very little if it's six. In the chart below, share of voice on ChatGPT and Gemini climbs from 17.2% to 23.1% over 14 days, a 5.9-point move well outside an 8% margin of error on a 12-prompt project, so it reads as a real trend, not noise. If you want to see this discipline applied under real pressure, the crisis communications playbook covers the same monitoring loop on a compressed timeline.
From snapshot to system
A baseline report is a fine pitch and a bad way to run an engagement. The value in AI visibility work shows up over weeks, in a trend line that separates real movement from the noise built into how AI generates answers. Get the competitor list right, get prompt coverage right, and put it all on a nightly cadence with someone checking the trend, and the report writes itself every month instead of becoming a scramble the day before a client call.
Once the tracking system is running, the next question is usually whether visibility is actually sending traffic. AI Traffic Analytics picks up there, measuring real AI-driven visits instead of estimating them from the outside.