Documentation

AI Search Citation Optimizer Methodology

What gets scored

The Citation Optimizer scores one page against a set of searches on up to four AI assistants: ChatGPT, Google AI Overviews, Google AI Mode, and Claude. The page can be a live URL, a pasted draft, or an earned-media placement.

The search set is one main search plus the related searches that go with it. AI assistants do not run your prompt as written; they write searches of their own and merge what comes back, so the page is scored across the whole set rather than one keyword. Each search is labeled with where it came from: observed in your tracked data, entered by you, or suggested by us. A search aimed at a single domain is kept as evidence and left out of the score, because no page can rank for it.

The four assistants find and read pages differently, so each runs its own checks. A check is a pass, warn, or fail decision, with the evidence behind it and a recommendation when it is not passing.

The checks each assistant runs

Every check carries a label for what can move it.

LabelMeans
RewriteThe words on the page can change it.
MetadataThe title, description, dates, or snippet directives.
TechnicalHow the page is served: rendering, robots rules, source order.
Off-pageNothing on the page moves it.
Not scoredReported for context, left out of the number.

Checks marked Not scored are excluded on purpose: scoring something the author cannot change would build an unfixable penalty into the number.

ChatGPT

CheckWhat it looks atWhat can move it
Triggers a searchWhether this search sends ChatGPT to the web at all.Not scored
Research routeWhether it routes to the cited-answer pipeline rather than shopping or news.Not scored
Named in searchesHow often ChatGPT writes your brand or domain into its own searches.Not scored
AccessWhether ChatGPT can open the page and read its text. It obeys robots.txt, does not run your JavaScript, and honors a request not to quote.Technical
Search coverageHow many searches in the set the page ranks for, and how highly.Off-page
Search listingThe title and description that have to earn the open.Metadata
Read budgetChatGPT reads a limited stretch from the top of the page, in the order your server sends it, then stops. Whether your best answer sits inside it.Rewrite
RelevanceYour strongest passage against the strongest passage on each ranking page, by exact wording, by meaning, and by a relevance model.Rewrite
SelectionWhether anything gets the page passed over once found, such as a date that reads as out of date.Rewrite
QuotabilityWhether the passage could be lifted into an answer as it stands: self-contained, answer-first, specific.Rewrite

Named in searches is the strongest signal we can see, and no rewrite moves it, so it is reported next to the score and kept out of it. Access is a hard check: if ChatGPT cannot read the page, the run stops and the later checks report as not evaluated rather than failed.

Google AI Overviews

Scored against the main search, because AI Overviews answers from a single retrieval.

CheckWhat it looks atWhat can move it
Can Google use itWhether the page is indexed, reachable by Googlebot, and allowed to show text from itself. noindex, nosnippet, data-nosnippet and max-snippet all gate it.Technical
Found in Google resultsWhether the page turns up in the results the overview is built from.Off-page
AI Overview todayWhether an overview fired for this search when we looked, and which sources it used.Not scored
What Google would quoteGoogle lifts individual sentences and joins them. This ranks your sentences and assembles the answer Google would build.Rewrite
Sentences stand aloneWhether a lifted sentence still makes sense without the page around it.Rewrite
Adds something newWhether the passage adds to what the other sources already say.Rewrite
Ready to be quotedThe shared writing check, applied to the passage Google would assemble.Rewrite

Google AI Mode

The same checks as AI Overviews, with two differences, because it is a separate retrieval pipeline rather than another view of the same one. It is scored against the whole search set, since AI Mode fans out on its own and answers across those sub-searches. And it has no AI Overview check, because it runs server-side with nothing on a results page to look at.

The pages the two use overlap very little, which is why they carry two scores rather than one.

Claude

Claude takes its own path. It searches Brave rather than Google, receives a short extract of each page rather than the page, filters results with code it writes, and quotes in very short spans.

CheckWhat it looks atWhat can move it
Would Claude searchWhether Claude would go to the web for this question at all.Not scored
Found by Claude's searchesHow many searches in your set return the page in Brave's results.Off-page
Inside what Claude readsWhether your strongest passage falls inside the short extract Claude receives.Rewrite
Exact words presentClaude's filter is code, and code matches strings rather than meanings. Whether the distinctive words, product names and figures from your searches are on the page.Rewrite
Short enough to quoteWhether your key sentences fit the short span Claude quotes, and name their own subject.Rewrite
Claude's history with your siteHow often Claude has used your site as a source in the answers we track.Not scored
Ready to be quotedThe shared writing check, applied to very short spans.Rewrite

How the combined score is formed

Each assistant scores itself first, as a weighted average over its own scored checks: a pass counts 100, a warn 50, a fail 0.

The combined score is a weighted average across the assistants. Weights are equal by default and you can change the mix on the results screen; the score, the range and the verdict recompute together from results already collected.

An assistant that did not finish is dropped from the average, out of the total as well as the top of it. It is never counted as zero, because a run that failed says nothing about the page. There is no floor on the combined score either: a page that does well on three assistants and fails outright on the fourth is reported as a good score with a warning beside it.

Before-and-after comparisons use a content-only sub-score, computed the same way over the rewrite-lever checks alone. It is the only figure that stays comparable between a published page, which has a search ranking, and a draft, which does not.

The range

Every score is shown as a range rather than a single number, because these pipelines do not make the same decision every time they see the same page. The range never narrows past a fixed minimum, and it widens when the judgment-based checks disagree between their two readings.

A move inside the range is labeled no meaningful change, in a neutral tone. It is not colored as an improvement and it cannot produce a ready-to-publish verdict. If a rewrite moved the score by less than the range, nothing has been shown to move.

How recommendations are merged

Four assistants produce four lists. Every finding maps to a shared signal, and the lists merge into one, in three kinds:

  • Agreed. Two or more assistants asked for the same thing. These come first, because they pay off everywhere.
  • One assistant. Only one asked. Worth doing, weighted by how much that assistant matters to your audience.
  • Trade-off. Two asks that cannot both be satisfied in the same span of text.

A trade-off is not resolved by averaging; averaging two incompatible instructions produces a rewrite that satisfies neither. It is resolved by giving each ask a different part of the page. When the exact question and the longer, more specific wording want the same spot, the question goes in the heading and the specifics go in the paragraph under it. The rewrite receives those allocations directly, so the draft honors both sides rather than the last instruction it read.

What a rewrite changes

It will front-load the answer, tighten passages into self-contained units, carry the specifics and exact terms into the part each assistant reads, split sentences that are too long to quote, move the answer inside the read budget, and produce a revised title, meta description and structured data. It comes with a change log tracing every edit back to the finding behind it.

It will not invent facts, statistics, credentials, customers or dates, and it will not add "best", "leading" or similar claims the page cannot support. The rewrite is meaning-preserving. If the page does not contain the number a check asked for, it tells you to add it rather than making one up.

Off-page findings are never rewritten. They appear in the list as context, clearly marked, and sit outside the rewrite's objective. A rewrite cannot change what Brave returns or whether ChatGPT names your brand.

Readiness verdicts

After every score the optimizer says what to do next. The verdict follows the current weights, so changing the mix recomputes it along with the score.

VerdictMeans
Revise againThere is headroom. At least one assistant has a selection check that is not passing.
Ready to publishEvery assistant that finished passes its selection checks and the combined score clears the publish bar.
PlateauedThe last pass did not move the score beyond its range. Use the best version so far.
Held backThe rewrite made at least one assistant meaningfully worse, even though the combined score held up.
PendingNothing has finished scoring yet.

A rewrite is supposed to raise the combined score without making any assistant worse, so a version where Claude fell while the average rose is held back and the previous version offered instead. Nothing is lost: the live page is untouched and the draft is still there if you disagree.

What the outline is built from

Scoring answers why a page was not cited. The outline answers what a page should say before it exists.

From a keyword and a page type, the outline resolves the searches to write for, reads what the ranking pages cover and what none of them answers, and fills a page-type template into a section-by-section brief. Each section carries its target searches, a word budget, the terms that have to appear, and the rules that decide whether a passage can be quoted. Score the draft when it is written; it runs against the same searches the outline was built from.

Scoring versions

Every run records the scoring version that produced it, shown next to the score. Scores from different versions are not comparable, and nothing in the product compares them. A run keeps rendering with the checks that produced it, and when a check or threshold changes materially the version increments instead of old runs being relabelled.

Limits

All of this approximates systems we do not have the source code for. The real rerankers are proprietary; we use an open-source model that behaves comparably, so the ordering it produces is the useful output rather than the absolute number. Search results move daily, and Google decides per search whether to show an AI Overview, so an absent overview means the check was skipped, not failed. We fetch pages ourselves as SpyglassesBot rather than as ChatGPT, Googlebot or Claude; we check robots rules for the real crawler names, but a site that treats those crawlers differently will look different to us than it does to them. Extract sizes and snippet lengths measure systems that change without notice, and thresholds and weights are starting values that get recalibrated as real runs accumulate.

Clearing every check does not guarantee a citation. It means the factors we can see and measure are in place. The answers themselves stay nondeterministic.

On this page

How-to guides for setting up Spyglasses to track your AI traffic