How to See Which Sites AI Trusts (and Fix the Ones It Gets Wrong)

Jim Wrubel

Jim Wrubel

9/9/2026

#PR#How-to#Workflows#Citations#AI Search Visibility
How to See Which Sites AI Trusts (and Fix the Ones It Gets Wrong)

Seeing which sites AI trusts means reading the site-scoped searches assistants run before they answer, across ChatGPT, Gemini, Claude, Perplexity, and Google AI Overviews. When a model searches site:acme.com pricing instead of searching the open web, it has already decided where the answer lives. The domain it names is one it treats as a source. It takes seven steps.

Start by reading the weekly consultation counts; your own site against competitors, plus the share of searching answers that consulted anyone at all. Treat the third-party rows as publisher intelligence, because a site the model reads on purpose is a better pitch target than one with a big domain authority score. Then check the flagged look-alike domains, the addresses that resemble yours but aren't in your account. A live site there means your authority is landing on somebody else's page; an empty one means the model invented an address and the answer went nowhere.

From there it's correction work. Claim the domains that really are yours, which moves their history into your own numbers, and dismiss the rest. Fix the canonical address in your Organization schema, directories, social profiles, and boilerplate, then note the date so the change in the data has a cause attached. Add recipients to the wrong-domain alert list, and schedule the weekly read instead of remembering it. Most brands see wrong-domain sightings drop to zero within six to ten weeks of correcting the source.

Here's how it usually surfaces. Somebody on the comms team is testing prompts before a launch, notices the assistant quoting a support policy nobody recognizes, and traces it back to a domain the company let lapse in 2021.

Why the site AI picks tells you more than the pages it cites

Most AI visibility reporting counts citations. A citation means one of your pages won a search and made it into the answer. That's worth counting, and it's not the same thing as trust.

A site-scoped search is a different move. The model isn't asking the web who has the answer; it's asking one particular site, because it already believes that site is where the answer lives. Nothing can rank for site:acme.com, so this is worthless as a ranking opportunity. As evidence it's about as direct as this channel gets.

That gives you three readings, and they mean different things.

The model consulted your site. It treats you as a source worth reading before it commits to an answer. This is the good case, and it tends to show up on consideration and decision questions before it shows up anywhere else.

It consulted a competitor and never you. The finding isn't that your number is low. It's that the model has a site it trusts for these questions and yours isn't it.

It consulted a third party. A publisher, a reference work, or a forum it trusts more than either of you. That's a PR input rather than a competitive one, and it's the most actionable row on the table.

Then there's the fourth case, which is the one that makes people sit up. The model consulted a domain that looks like yours and isn't. Your authority is real enough that it went looking for you by name; it just went to the wrong address. This workflow works with any AI visibility setup, and the callouts show how each step runs in Spyglasses.

StepWhat you're doingWhat tells you it's working
1. Read the consultationsCounting who AI reads directlyA weekly number for your own site
2. Rank the third partiesTurning trust into a pitch listTargets picked by evidence, not authority scores
3. Check the look-alikesFinding out what's at each addressEvery flagged domain explained
4. Claim or dismissSorting yours from noiseHistory reclassified, alerts quiet
5. Fix the sourceOne canonical address everywhereA dated annotation on the timeline
6. Turn on alertsRouting recurrences to emailA named recipient list
7. Automate the readScheduling the weekly summaryA digest that arrives without a reminder

Read which sites AI consulted before it answered

Start with the weekly counts, and read three numbers rather than one.

Your own site first. That's the count of times an assistant ran a scoped search against your domain in the last seven days. On its own it's a raw number, and raw counts mostly measure how chatty a platform was that week. One prompt can fan out into a dozen searches.

So read the rate next. Of the answers where the model searched the web at all, what share included a consultation of you? That question has a stable answer week to week, which makes it the number worth putting in a report.

Third, the split. Your site against competitors' against third parties. A brand that gets consulted twice while a competitor gets consulted thirty times has a specific problem, and it isn't a content volume problem.

One catch worth knowing before you quote any of this to a client. Prompts that name your brand outright, and head-to-head comparison prompts, should be held out of the headline numbers. A question that names you is expected to send the model to your site, so counting it lets the metric congratulate itself. A brand tracking mostly branded prompts would look maximally authoritative while telling you nothing. If your setup doesn't exclude them, exclude them by hand before you report.

Turn the sites AI reads directly into pitch targets

This is the step PR teams get the most out of, and it's the one people skip because the row looks like background noise.

A third-party domain in that table is a site an assistant chose to read before answering a question in your category. Not a site that happened to rank. A site it went to on purpose. That's a much stronger signal than domain authority, which measures how the open web links to a publication rather than whether a model reads it.

So rank your outreach by it. A trade publication with modest traffic that gets consulted directly on eleven answers a week is a better target than a famous outlet that never appears. Reference works and well-moderated forums show up here more than most comms teams expect, and they're often easier to correct than a magazine feature.

Two things decide whether a target is worth the pitch. Whether AI can crawl the site at all, and whether it already shows up for questions in your category. A well-known publication that blocks AI crawlers in its robots.txt won't move an answer no matter who you know there. Scoring the list before you work it is the same discipline as any earned media program in this channel; connecting earned media to AI visibility covers the scoring in depth, and building a PR pitch list AI can actually see covers turning it into outreach.

Investigate the domains that look like yours but aren't

Now the interesting part. Some flagged rows are addresses that resemble your brand and aren't in your account: a different top-level domain, a country variant, a hyphenated spelling, the name of a product you acquired.

Open each one in a browser. What you find sorts into four cases, and each has a different fix.

What's at the addressWhat it meansWhat to do
A live site that is really yoursA legacy, regional, or acquired domain missing from your accountClaim it, then set the canonical address
A live site somebody else runsYour authority is landing on their pageFix the source, and consider outreach
A parked page or domain-for-sale listingThe model trusts an address with nothing behind itFix the source; buy it back if it's cheap
Nothing served at allThe model invented the address; the answer went nowhereFix the source, and check what it cited instead

The first case is the most common and the least alarming. Companies accumulate domains. A regional site from a market you entered in 2019, a top-level domain you bought defensively, the address of a company you acquired two years ago. The model isn't wrong that those are you; your account just doesn't know it yet.

The last case is the one worth pausing on. When nothing is served, the assistant made up a plausible address from your brand name, searched it, found nothing, and answered anyway using whatever else it had. Nobody saw an error. The answer just got built from weaker material.

One sighting isn't a pattern. What matters is recurrence, the same brand-like domain turning up across separate runs inside a week. A single odd search is usually a one-off; the same wrong address three times means the model learned it somewhere and kept it.

Pull quote: A site AI reads before it answers is worth more than a site it cites in passing. One is a source it trusts. The other is a link it found.
A site AI reads before it answers is worth more than a site it cites in passing. One is a source it trusts. The other is a link it found.Spyglasses

Claim or dismiss each wrong address

Every flagged domain should end up in one of two states, and leaving them unsorted is what turns a useful alert into background noise people stop reading.

Claim the ones that are yours. That means the legacy top-level domains, the regional sites, the domains that came with an acquisition, and anything else your company actually controls. Claiming reclassifies the history too, so past consultations of that address move out of the third-party column and into your own numbers. Your authority reading gets more accurate going backward as well as forward, and the alerts for that domain stop.

Dismiss the rest. A dismissed domain stays counted as a third party everywhere it appears; you're only turning off alerting for it. Use this for the parked pages, the unrelated companies with similar names, and the addresses that don't exist. You want the alert list to mean something, and it only means something when the known cases are cleared.

Keep a short note on why you dismissed each one. Six months later, when the same domain resurfaces in a different context, the note saves somebody the same twenty minutes of investigation.

Fix the address everywhere the model learned it

Claiming a domain corrects your reporting. It doesn't correct the model. For that you have to fix the address at the places the model read it, and then wait for a recrawl.

Work through the durable sources in rough order of how much weight they carry.

Your structured data. One canonical url in your Organization schema, with every legitimate variant listed under sameAs. This is the single most direct statement you can make about which address is you, and plenty of sites still carry a stale one from a template set up years ago.

Directory and database listings. Crunchbase, industry directories, chamber listings, review platforms. These update slowly and get quoted with more confidence than they deserve, because a retrieval system reads three sites saying the same wrong thing as corroboration.

Social profiles. The link in every bio, on every platform, including the accounts nobody has posted to since 2022.

Press boilerplate. The "About" paragraph at the bottom of every release you've issued. It gets copied verbatim into coverage, which means one wrong URL in boilerplate propagates into dozens of pages you don't control.

Redirects. If a domain is yours and no longer serves the main site, make it redirect rather than sit parked. A permanent redirect gives the crawler an unambiguous answer.

Then note the date you made the change. The lag between fixing an address and seeing the wrong-domain sightings stop runs six to ten weeks, sometimes longer for directory data, and nobody remembers in November what changed in September. If the fix is part of a wider name or domain change, the sequencing in tracking a rebrand in AI answers applies here too. It's also worth confirming assistants can read the canonical site cleanly; the AI readiness audit shows which pages parse without trouble.

Turn on alerts so a repeat reaches an inbox

Detection runs whether or not anyone is watching, and that's exactly the problem. A finding that only lives on a timeline gets seen when somebody opens the timeline.

Add recipients to the wrong-domain alert list in your property settings. Two or three people is right; whoever owns the website, whoever owns comms, and whoever would field the question if a customer sent a screenshot. An empty list means the detection still runs and the finding waits for a visit.

This matters most in the weeks nobody is watching closely. Wrong addresses tend to appear right after the events that create them, which are the same events that keep everybody busy: a domain migration, a rebrand, an acquisition, a regional site launch. Turning the list on before that work starts is worth the two minutes.

Automate the weekly read

The last step is making sure this actually happens every week rather than the week after somebody remembers it exists.

The numbers are small enough to summarize in a paragraph, which makes them a good fit for a scheduled digest. Pull the current seven days and the seven before it, and narrate three things: your own consultations against the previous window, any new third-party domain that appeared, and any flagged look-alike. If none of those changed, the summary says so in a sentence and you move on.

Connect your AI assistant to your visibility data and the digest writes itself. A weekly scheduled prompt that reads the last seven days against the previous seven costs nothing to run and turns a dashboard visit into an email you skim. That's the difference between a report somebody checks in busy weeks and one they check every week.

Keep a human in the loop for the interesting rows. The digest is good at telling you a new domain showed up. Deciding whether it's a pitch target, a claim, or a dismissal is still a judgment call, and it's the one part of this workflow that stays manual on purpose.

What good looks like after a quarter

This workflow pays off in small, specific ways rather than one big chart. Three months in, you should have four things.

  1. A weekly number for your own site. Consultations of your domain, and the share of searching answers that included one, tracked long enough to have a normal range. That range is what makes a bad week legible as a bad week.
  2. A pitch list built from evidence. The third-party sites AI reads directly in your category, scored and ranked. It's a better list than the one you'd build from domain authority, and it costs nothing extra to produce.
  3. A clean domain roster. Every look-alike address either claimed or dismissed, with the reasoning written down, and alert volume for the sorted ones at zero.
  4. One documented correction. The canonical address fixed at the source, the date noted, and the wrong-domain sightings falling off over the following weeks. One cause and effect you can point at beats a quarter of charts.

The habit is ten minutes a week on the consultations view. Most weeks nothing has changed, which is the correct outcome for a metric that runs on recrawl cycles. The weeks it does change, you find out that an assistant has been sending people to an address you don't own, and you find out before a customer does.

How to See Which Sites AI Trusts (and Fix the Ones It Gets Wrong)