Citation Optimizer
What this dashboard does
The Citation Optimizer takes one page, a set of searches, and runs them through ChatGPT, Google AI Overviews, Google AI Mode and Claude. It gives you one combined score with a range, each assistant's own score and checks, one merged list of changes, and a rewrite that acts on all four at once.
Where AI Visibility Rankings tells you which searches you show up for, the Citation Optimizer tells you why a specific page does or does not get cited, and what to change. For the model behind the checks, see the Citation Optimizer methodology.
The free plan scores ChatGPT only, against a single search. Paid plans score all four assistants against a full search set.
One-time setup: crawl and index your pages
Accurate page matching needs your citation-relevant pages crawled and indexed. A new property starts with none, so the first time you open the optimizer you will see a setup banner offering to crawl and index.
Click to start it. The crawl runs in the background and the banner shows progress; an embedding pass afterwards can take a minute or two. This is a one-time step per property. Once it is done, the optimizer can match a search to your closest existing page and detect overlap between your pages.
Product properties skip this step. Crawled page records only exist for a property with a crawlable site of its own, and a product's pages live on the parent company's domain. A product works from Paste a URL and Paste a draft instead. The optimizer needs a content anchor to score against, so it is not offered on person, category or location properties.
Setup uses no credits. Scoring uses no credits either; see What it costs.
Setup: choose the searches, the assistants, and the page
The search set
The first panel is Searches to score against. This is the scoring unit: one primary search plus the related searches you want the page to hold up across. Two to six is the useful range.
Every row carries a provenance badge, and the badge follows the search everywhere it goes afterwards:
| Badge | Means |
|---|---|
| Tracked | Observed. Your assistants were seen running this search, captured from your latest AI Visibility report. |
| Entered | You typed it. Scored identically, just not confirmed as a search any assistant runs. |
| Suggested, not observed | The optimizer proposed it. Nobody has seen an assistant run it for this brand. |
The picker lists your property's tracked searches first, highest impact first, with a flag on the ones where you do not rank yet. Those are usually the highest-value runs.
Site-scoped searches stay visible, marked "Used as evidence, not scored." A search the model aimed at one domain (site:example.com pricing) has no open-web result to win, so nothing can rank for it. It is kept in the set because it is real evidence about which sites the model consults directly, and it feeds the named-in-searches diagnostic. It cannot be the primary search. Hovering the badge explains why.
The assistants
Below the searches is a row of platform cards: ChatGPT, Google AI Overviews, Google AI Mode, Claude. Pick which ones to score against; by default you get every one your plan covers.
AI Overviews and AI Mode are separate cards on purpose. They are different retrieval pipelines that end up using surprisingly different pages, so they carry separate scores and separate detail pages.
Cards your plan does not cover show a lock with the reason and an upgrade link, rather than disappearing. On the free plan, three of the four are locked.
A card can also be temporarily unavailable if one of the pipelines is paused for an outage. That is an availability state, not a plan state, and it says so.
The usage meter
At the bottom, next to the button that spends from it, is the monthly page meter: pages used, pages allowed, and when it resets.
It counts pages, not runs. Re-running the same page against different searches or different assistants still counts once for the month, and re-scoring a rewrite never counts at all. At the cap, the score button is disabled and reads Monthly page limit reached, with the reset date and an upgrade path.
The page
Above everything is the page being scored. Three ways to supply it:
- Pick a page from your crawled and indexed pages, ranked by how well their content matches the primary search.
- Paste a URL, including an earned-media placement on a publisher you do not own.
- Paste a draft in markdown, for a page you are still writing.
Pick a page type (product, homepage, informational, press release, or general). It drives the rewrite template later.
Results: the combined score
The results screen is the whole run in one view.
The combined score with its range
The headline number is the weighted mean of the assistants that finished, and it is always shown with its range. The range is at least ±8 points wide, and wider when the judgment-based checks disagreed with themselves.
Treat the range as the unit, not the number. A page that moved from 54 to 59 has not moved.
Platform tiles
Four tiles, one per assistant, each with its own score, its own range, its verdict in a phrase, and a link into its detail page.
If an assistant did not finish, an alert above the score says so by name. That assistant is left out of the weighting rather than counted as zero, so the combined score reads as "the score of the assistants we could actually measure". The failed tile carries its own Try again action, and re-scoring after a rewrite retries it automatically.
Weighting
Under the combined score, behind a disclosure, is the weighting control: presets plus a slider per assistant, equal by default, with Reset to equal.
Re-weighting recomputes the score, the range and the readiness verdict from results already collected. It costs nothing, uses no page from your allowance, and updates instantly. Use it when you know where your audience actually is. A mix that leans on the two assistants with the widest ranges will visibly widen the combined range, which is the correct outcome rather than a bug.
The mix is saved on the run, so a shared link shows the same numbers you saw, and it becomes the property's default for its next run.
Named in searches
Below the score, in a dashed callout, is the named-in-searches diagnostic: how often the assistants wrote your brand or domain into the searches they ran for themselves.
It is the strongest known predictor of being cited, and it is not part of the score, because no rewrite can change it. The callout says so. If this is low, the work is brand demand and earned coverage, not paragraphs.
The coverage table
The last panel is a grid of searches against assistants: for each search in your set, which assistants return the page and roughly where. It is the fastest way to see whether you have a writing problem or a coverage problem. A page that only ever comes back for one search in the set has a coverage problem, and the checks below will keep telling you the same thing in four different ways.
Platform detail
Each assistant has its own page at its own URL, so a failing check can be pasted into a ticket. The switcher between assistants is links, not tabs, for the same reason. A URL can also open straight onto a named check.
Every detail page has the same shape:
- The assistant's score with its range, and the scoring version that produced it (for example, "ChatGPT scoring v2").
- A step strip: every check in pipeline order, colored pass, warn, fail, or not evaluated, with the selected one highlighted.
- The selected check, expanded: what it asks, what it found, its lever chip (Content, Metadata, Technical, Off-page) and its recommendations.
The lever chip is there so you can tell, before spending anything on a rewrite, whether a rewrite is the right tool. A page failing on search ranking and site familiarity is not a page a rewrite fixes, and the chips say so next to the score.
The featured panels
One check on each assistant opens into raw evidence rather than a summary. These stay on the detail page rather than routing away, because each one is meaningless without the check it belongs to.
ChatGPT
- Read budget. Your page laid out as ChatGPT receives it, with the roughly six-thousand-character line drawn across it and your best answer passage marked either inside or outside the window. This is the panel that explains a low ChatGPT score more often than any other.
- Competitor passages. Your strongest passage side by side with the strongest passage from each page that ranks for the search, in merged rank order, so you can see what you are being compared against rather than being told a number.
Google AI Overviews
- Snippet preview. The answer Google would assemble from your page: your sentences ranked against the search, the ones it would lift, joined the way Google joins them. Headings are scored as sentences because Google treats them that way.
- AI Overview sources. If an overview fired for this search when we looked, the sources it used. If none fired, the panel says the check was skipped, not failed, because Google makes that decision per search and it changes day to day.
Google AI Mode
- Set coverage. AI Mode fans out on its own, so its featured panel is coverage across the search set rather than a snippet simulation.
Claude
- Exact wording found. The distinctive words, product names and figures from your searches, each marked present or missing on the page. Claude filters its results with code, and code matches strings rather than meanings, so this panel is often the difference between a good Claude score and a poor one.
- Quotable sentences. Your most relevant sentences with the 150-character limit drawn on them, and a note on whether each one names its own subject. Anything longer is not quotable by Claude at all.
Recommendations
The recommendations screen merges all four assistants into one list, in three kinds, in this order:
Trade-offs first. Two asks that cannot both be satisfied in the same span of text. These come first because they change how every other edit gets made. Each one shows both asks in the words of the assistant that made them, plus a resolution that gives each ask a different part of the page. A trade-off is never resolved by averaging; averaging two incompatible instructions produces a rewrite that satisfies neither.
Consensus next. Two or more assistants asked for the same thing. Do these first among the ordinary items; they pay off everywhere.
Platform-specific last. One assistant asked alone. Worth doing, weighted by how much that assistant matters to you.
Every item carries its severity, the assistants that raised it, the part of the page it applies to, and its lever. Off-page items are in the list as context and are clearly marked as things a rewrite will not touch.
Revise and re-score
Click Revise this page. The rewrite is grounded in every assistant's findings at once, with the trade-offs resolved by allocation, and it produces a revised draft plus a revised title, meta description and page-appropriate structured data. It never invents facts, statistics or credentials, and it will not add superlatives the page cannot support.
The revision view shows:
- Before and after, with both ranges. Each assistant's move is compared against its own range before it is given a color.
- "No meaningful change" in a neutral tone wherever a move is inside the range. It is not a small green delta, because it is not an improvement.
- A change log tracing each edit back to the finding that motivated it.
Then re-score. Re-scoring runs the same search set against the same pinned competitor pages, which is what makes the before and after mean anything: the rewrite is measured against identical competitors rather than against whatever happens to rank today. Re-scoring is always free and never counts against your page allowance.
The readiness verdict
A banner tells you what to do next. It follows the current weights, so re-weighting updates it along with the score.
| Verdict | What to do |
|---|---|
| Revise again | There is headroom. At least one assistant has a selection check that is not passing. |
| Ready to publish | Every assistant that finished passes its selection checks and the combined score clears the bar. Publish it. |
| Plateaued | The last pass did not move the score beyond its range. Use the best version so far. |
| Held back | The rewrite made at least one assistant meaningfully worse, even though the combined score held up. The previous version is offered instead, by name. |
Held back is worth understanding. The rewrite's objective is to raise the combined score without making any assistant worse, so a version where Claude fell 12 points while the average rose has not met it. The combined number alone would not have told you.
Held back is not a destructive state and does not render as one. Nothing is lost: your live page is untouched, the draft is still there, and you can use it anyway if you disagree with the call.
Outlines: before the page exists
Scoring answers "why was this page not cited". An outline answers "what should the page say in the first place".
Start from a single keyword and a page type (homepage, product, informational or press release; each has its own template). The optimizer resolves the searches to write for, reads what the pages that already rank cover and what none of them answers, and produces a writer brief:
- Section by section, with the target searches for each section.
- Word budgets and must-include terms per section.
- An FAQ built from the searches nobody is answering well.
- The coverage gap: the thing none of the ranking pages covers.
- The rules that decide whether a passage can be quoted, written as instructions a writer can follow.
- Revised meta title and description to write toward.
Export it as markdown or HTML; both paste cleanly into a document.
Provenance survives into the brief. Every search in it is labeled tracked, entered, or suggested-and-not-observed. Keep the labels when you hand the brief to a writer: by that point, nothing else tells a suggestion apart from an observation.
Once the draft is written, score it from the brief. It scores against the same searches the brief was built from, so the two screens tell one story.
Outlines live in their own list. They belong to a property, not to a scoring run.
Driving it from an AI assistant (MCP)
The whole loop is available to a connected assistant through the Spyglasses MCP server: pick the searches, find the page, score it across every assistant your plan covers, read the merged recommendations and the readiness verdict, generate a rewrite, re-score, and stop when the verdict says to.
See MCP: Citation Optimizer tools for the full tool list and the fire-and-poll loop, and Connecting the MCP server for setup. Everything is scoped to your own properties and nothing is published on your behalf.
What it costs
- Scoring costs no credits, on every plan. It is metered by pages instead: each plan includes a number of pages a month.
- A page counts once a month, however many times you re-run it with different searches or assistants.
- Re-scoring a rewrite never counts. Once you are in the loop on a page, the loop is free.
- Re-weighting never counts. It recomputes from results already collected.
- Rewrites and outlines cost credits. Paid plans include a monthly allowance of rewrites, and outlines draw on the same allowance; beyond it, each is billed per use.
See the pricing page for current numbers, and Settings, Billing for your credit history, where these appear as "Citation Optimizer".
Overlapping pages
When the optimizer matches a search to your pages, it also flags overlapping pages: pairs of your own pages so similar that AI search cannot tell which one to cite. Overlap splits your authority, so instead of one strong page ranking, two weaker pages compete for the same citation and neither wins cleanly.
Two fixes:
- Consolidate the pages into one stronger page and redirect the weaker one.
- Differentiate their focus so each targets a distinct search.
How it is detected: overlap uses a dual signal. Two pages are flagged only when both their full-content embeddings and their meta-description embeddings are highly similar (each above roughly 0.92 cosine similarity). The dual check matters: a single-signal version flagged distinct-but-templated pages as overlapping. Requiring the meta description to corroborate the body content removed those false positives.
On heavily templated sites, shared layout can still pull two truly different pages closer together than their unique content warrants. Per-section page vectors would sharpen this further; that is on the roadmap.
Tips
- Build the set, do not settle for one search. Appearing across several related searches is the strongest content-adjacent lever on every assistant. A set of four or five related tracked searches will teach you more than one search scored four times.
- Read the levers before you buy a rewrite. If the failing checks are off-page, a rewrite will not fix them and the screen will tell you that for free.
- Start from a gap. The highest-value runs are tracked searches where you do not rank yet. The optimizer shows you which competitor pages to study and which of your pages is closest to competing.
- Score a draft before you publish it. Paste it in and check it against its target searches while it is still cheap to change.
- Do not chase moves inside the range. If the screen says no meaningful change, another automated pass is unlikely to help.
- Re-weight to match your audience, not to reach a number. The verdict follows the weights, so a mix chosen to manufacture "ready to publish" only fools you.
- Fix overlap first. Consolidating or differentiating two competing pages often does more than any single content edit.