How to Monitor What AI Says About Your Brand

Jim Wrubel

Jim Wrubel

9/6/2026

#In-house#How-to#Workflows#Brand Safety#AI Search Visibility
How to Monitor What AI Says About Your Brand

Monitoring what AI says about your brand means asking the assistants a fixed set of questions on a schedule and keeping the answers, because an AI answer isn't posted anywhere for a listening tool to catch. The setup has five parts. Write down the specific claims you can't afford to have repeated. Turn them into a small set of reputational questions a worried customer would actually ask. Run those questions nightly across ChatGPT, Gemini, Claude, Perplexity, and Google AI Overviews. Read how often each claim surfaces, week over week. Then pull the citations behind those answers to find the pages feeding the claim. A claim that shows up in 18% of answers this week and 11% next week is fading. One that holds steady is being fed by a page you can find and work on. Expect four to ten weeks between fixing the source material and seeing the claim drop out, because the page has to be recrawled and start winning retrieval before anything changes.

Here's the meeting that starts this workflow. Someone in legal or comms types the company name into an AI assistant, gets an answer with a line in it nobody approved, and asks the marketing team a simple question. How long has it been saying that, and how often?

Why the old monitoring stack misses this

Brand monitoring was built for content that sits still. A post, an article, a review; each one has an author, a date, and a URL. You can count them, alert on them, and screenshot them for the deck.

An AI answer has none of those properties. It gets generated for one person, in one session, and then it's gone. Nobody files it. Two people asking the same question can get different wording, and neither version exists anywhere you can search. So a claim can be in heavy rotation for months without producing a single alert in any listening tool you already pay for.

The other difference is reach. A bad review sits at the bottom of page three. An AI answer arrives as the answer, in one paragraph, with the assistant's confidence behind it. If a wrong claim about your safety record, your pricing, or your ownership is in that paragraph, it lands on every person who asks, and none of them see the source it came from unless they go looking.

That's the gap this workflow closes. You can't wait for mentions to arrive, so you ask the questions yourself, on a schedule, and keep the record. It works with any AI visibility setup; the callouts show how each step runs in Spyglasses.

StepWhat you're doingWhat tells you it's working
1. List the claimsNaming the exact statements you're watching forA short list legal signed off on
2. Build the questionsWriting what a worried buyer would askQuestions that pull real answers, not ads
3. Run them nightlyPutting the set on a daily cycleA trend line instead of a screenshot
4. Read the surfacing rateTracking each claim over timeMovement you can point at a cause
5. Trace the sourcesFinding the pages behind the claimA short outreach list, ranked

List the claims you can't afford to have repeated

Start with the list, not the tooling. Get comms, legal, and whoever owns the product story in one room and write down the specific statements that would cause a problem if an assistant repeated them.

Be specific. "Negative sentiment" isn't something you can track. "The company was fined for the 2024 data breach" is. So is "their contracts are hard to cancel," "the product is discontinued," and "they were acquired last year," if that last one isn't true. The more exact the claim, the more useful everything downstream gets.

Write each one the way a customer would say it, not the way your lawyers would. Assistants generate plain language, so a claim phrased in legal terms will slip past a match that a plain-English version catches. If your matching is semantic rather than keyword based, a paraphrase still counts, which is what you want here; the same idea comes back in different words every night.

Then write the counter-claims. These are the true statements you want showing up instead; the certification you hold, the policy you changed, the number that's actually correct. You're going to measure both. One list should be falling and the other should be rising, and having them side by side is what turns this from monitoring into a scoreboard.

Sort the list before you move on. Not everything on it deserves the same response.

Claim typeWhy it lands in AI answersHow fast to respond
Factually wrong (dates, ownership, status)An old page still ranks and nothing newer corrects itFast; publish the correction now
True but outdated (a fixed problem)The coverage of the problem outlives the fixWeeks; get the resolution documented
True and currentThe source material is accurateNot a monitoring problem; it's a business one
Competitor framingA comparison page tells the story their wayWeeks; answer the comparison yourself
Category risk (an industry claim)The assistant fills a gap with generic industry textSlow; it moves when your own pages get clearer

Ask the questions a worried customer would ask

Now write the questions. This is the step people rush, and it's the one that decides whether the whole setup tells you anything.

The instinct is to ask about the claim directly. "Was [brand] fined for a data breach?" That question does need to be in the set, but on its own it's misleading, because you asked a leading question and the assistant went looking for material about a fine. Of course it found some.

The questions that matter are the ones a real person asks when they aren't thinking about your problem at all.

  • Decision questions. "Is [brand] safe to use for healthcare data?" "Should I sign a contract with [brand]?" These are where a harmful claim does actual damage, because it surfaces in front of someone about to buy.
  • Category questions. "What's the most reliable option for [category]?" If your claim shows up here, unprompted, that's the highest-severity version of the problem.
  • Direct questions. "What problems have people had with [brand]?" The leading version. Keep a few; they tell you what material exists, which is useful even when the claim isn't surfacing anywhere else.
  • Comparison questions. "[Brand] vs [competitor]." Assistants love these, and a rival's comparison page is often where a claim about you gets its cleanest phrasing.

Ten to twenty questions covers most brands. Resist the urge to write fifty. Every tracked question runs on a schedule and costs on its own, and a set nobody reads every week is worse than a small one somebody does.

Tag the whole set. One tag, something like compliance or brand safety, applied to every question in the group. That tag is what lets you filter every chart and export down to this view later without rebuilding the list from memory.

Put the questions on a nightly schedule

A spot check tells you almost nothing. Ask an assistant the same question three times and you can get three different answers, and any one of them could be the outlier. What you need is a rate, and a rate needs repetition.

So put the tagged set on a nightly cycle across the major assistants and let it collect. Within two weeks you have a baseline. Within six you have a trend, and the trend is the thing you'll actually report on.

Two setup details are worth getting right on day one.

Run every assistant, not only the one your team uses. They retrieve different sources and land on different claims, and it's common for a claim to be heavy in one and absent in another. That split is a useful signal by itself; it usually means one specific source is carrying it.

Check the margin of error before you promise anyone a number. Fourteen questions running nightly gives you a directional read, which is enough to say a claim is rising or falling. It's not enough to say it moved 2 points. Say the plain version to your stakeholders early and you won't have to walk anything back later.

If you haven't set up general visibility tracking yet, do that first; the monitoring set works better as a filtered view of an established baseline than as a standalone project. Baselining a brand's AI visibility covers that setup end to end.

Pull quote: AI doesn't repeat what's true. It repeats what it can find, and most of what it can find about your brand was written by somebody else.
AI doesn't repeat what's true. It repeats what it can find, and most of what it can find about your brand was written by somebody else.Spyglasses

Watch how often the risky claims show up

Once the data is coming in, the discipline is to read the rate and ignore the individual answer. Someone on your team will find one bad response and forward it around. The response to that is to look up what percentage of answers carried the claim this week, and whether that percentage is going up or down.

Set a weekly rhythm. Fifteen minutes, filtered to the tag, looking at four things.

Surfacing rate per claim. How often each claim appeared, as a share of the answers. This is your headline number, and it's the one that goes to legal.

Direction over four weeks. One week is noise. A month of decline is a result, and a month of climbing is a problem worth escalating before it gets bigger.

Which assistant carries it. A claim concentrated in one assistant is usually one source. A claim spread evenly across all of them is in the general material about your category, and that takes longer to move.

Counter-message pull-through. Are your true statements showing up? A claim that's falling while nothing replaces it means the assistant just stopped mentioning the topic. A claim that's falling while your correction rises means the work landed.

Mark the events on the timeline as you go. A statement, a policy change, a piece of coverage, a page you published. Six weeks later, when the line moves, you'll want to know what happened that week, and nobody remembers.

One escalation rule is worth writing down in advance. Decide, now, what surfacing rate on a decision question triggers a bigger response, and what that response is. If a harmful claim hits your threshold, this monitoring workflow hands off to a faster one; running crisis communications when AI is repeating the story picks up from there.

Trace the sources feeding the answers

This is the step that turns monitoring into something you can act on, and it's the one most teams never get to.

When an assistant makes a claim about you, it built that claim from pages it retrieved. Those pages are visible. Pull the citations from the answers where the claim appeared, and you get a list of the material actually driving it. Almost always it's shorter than people expect. Three to eight sources carry most of a persistent claim.

The list usually sorts into four kinds of source.

A single old article. One piece of coverage from years ago that still ranks for your brand and hasn't been updated since the situation changed. The most common cause of a stale claim, and the most fixable.

An aggregator or a directory. Sites that scraped a fact once and never refreshed it. Low authority individually, but there are often several saying the same wrong thing, which reads as corroboration to a retrieval system.

A competitor's comparison page. Written to be retrieved for exactly the question you're worried about, and doing its job well.

Your own site. This one surprises people. An outdated policy page, an old FAQ, a support doc nobody archived. If your own material is feeding the claim, that's the fastest fix available, and it's worth checking your pages are readable to assistants at all; the AI readiness audit shows which of your pages an assistant can actually parse.

Score the outside sources before you spend outreach time on them. Two things decide whether working a source is worth it; whether AI can read that site at all, and whether it already gets cited for questions in your category. Plenty of well-known publications block AI crawlers, and a correction there won't change an AI answer no matter who signs it. Chase the sites the assistants are actually reading.

What good looks like after a quarter

Set expectations before the first report goes to legal, because the first month produces a baseline rather than a win, and that's the correct outcome.

By the end of a quarter, three things should exist.

  1. A defensible number per claim. Not a screenshot of one bad answer, but a surfacing rate with a trend behind it and a stated margin of error. That's what makes the topic discussable at a senior level instead of alarming.
  2. A source list you keep. The pages feeding each claim, scored and sorted. This doubles as the outreach and content plan for the next quarter, and it costs nothing extra to produce.
  3. One closed loop. Pick the most fixable claim, correct the source material, and watch the rate come down. One documented cause and effect is worth more internally than a year of monitoring with nothing attached to it.

The habit that keeps this working is small. Fifteen minutes a week, the same filtered view, the same four numbers. Most weeks nothing will have changed, which is the point. You're building the record that lets you answer the question at the top of this article with a date and a number instead of a guess.

And add new claims as they come up. A product recall, a leadership change, a lawsuit, a rumor a sales rep keeps hearing on calls. Each one takes about five minutes to add to the list, and the earlier it goes in, the more of its history you'll have when someone finally asks about it.

How to Monitor What AI Says About Your Brand