AI Search Citation Optimizer Methodology
What gets scored
The Citation Optimizer scores one page against a set of searches on up to four AI assistants: ChatGPT, Google AI Overviews, Google AI Mode, and Claude. The page can be a live URL, a pasted draft, or an earned-media placement.
The search set is one main search plus the related searches that go with it. AI assistants do not run your prompt as written; they write searches of their own and merge what comes back, so the page is scored across the whole set rather than one keyword. Each search is labeled with where it came from: observed in your tracked data, entered by you, or suggested by us. A search aimed at a single domain is kept as evidence and left out of the score, because no page can rank for it.
The four assistants find and read pages differently, so each runs its own checks. A check is a pass, warn, or fail decision, with the evidence behind it and a recommendation when it is not passing.
The checks each assistant runs
Every check carries a label for what can move it.
| Label | Means |
|---|---|
| Rewrite | The words on the page can change it. |
| Metadata | The title, description, dates, or snippet directives. |
| Technical | How the page is served: rendering, robots rules, source order. |
| Off-page | Nothing on the page moves it. |
| Not scored | Reported for context, left out of the number. |
Checks marked Not scored are excluded on purpose: scoring something the author cannot change would build an unfixable penalty into the number.
ChatGPT
| Check | What it looks at | What can move it |
|---|---|---|
| Triggers a search | Whether this search sends ChatGPT to the web at all. | Not scored |
| Research route | Whether it routes to the cited-answer pipeline rather than shopping or news. | Not scored |
| Named in searches | How often ChatGPT writes your brand or domain into its own searches. | Not scored |
| Access | Whether ChatGPT can open the page and read its text. It obeys robots.txt, does not run your JavaScript, and honors a request not to quote. | Technical |
| Search coverage | How many searches in the set the page ranks for, and how highly. | Off-page |
| Search listing | The title and description that have to earn the open. | Metadata |
| Read budget | ChatGPT reads a limited stretch from the top of the page, in the order your server sends it, then stops. Whether your best answer sits inside it. | Rewrite |
| Relevance | Your strongest passage against the strongest passage on each ranking page, by exact wording, by meaning, and by a relevance model. | Rewrite |
| Selection | Whether anything gets the page passed over once found, such as a date that reads as out of date. | Rewrite |
| Quotability | Whether the passage could be lifted into an answer as it stands: self-contained, answer-first, specific. | Rewrite |
Named in searches is the strongest signal we can see, and no rewrite moves it, so it is reported next to the score and kept out of it. Access is a hard check: if ChatGPT cannot read the page, the run stops and the later checks report as not evaluated rather than failed.
Google AI Overviews
Scored against the main search, because AI Overviews answers from a single retrieval.
| Check | What it looks at | What can move it |
|---|---|---|
| Can Google use it | Whether the page is indexed, reachable by Googlebot, and allowed to show text from itself. noindex, nosnippet, data-nosnippet and max-snippet all gate it. | Technical |
| Found in Google results | Whether the page turns up in the results the overview is built from. | Off-page |
| AI Overview today | Whether an overview fired for this search when we looked, and which sources it used. | Not scored |
| What Google would quote | Google lifts individual sentences and joins them. This ranks your sentences and assembles the answer Google would build. | Rewrite |
| Sentences stand alone | Whether a lifted sentence still makes sense without the page around it. | Rewrite |
| Adds something new | Whether the passage adds to what the other sources already say. | Rewrite |
| Ready to be quoted | The shared writing check, applied to the passage Google would assemble. | Rewrite |
Google AI Mode
The same checks as AI Overviews, with two differences, because it is a separate retrieval pipeline rather than another view of the same one. It is scored against the whole search set, since AI Mode fans out on its own and answers across those sub-searches. And it has no AI Overview check, because it runs server-side with nothing on a results page to look at.
The pages the two use overlap very little, which is why they carry two scores rather than one.
Claude
Claude takes its own path. It searches Brave rather than Google, receives a short extract of each page rather than the page, filters results with code it writes, and quotes in very short spans.
| Check | What it looks at | What can move it |
|---|---|---|
| Would Claude search | Whether Claude would go to the web for this question at all. | Not scored |
| Found by Claude's searches | How many searches in your set return the page in Brave's results. | Off-page |
| Inside what Claude reads | Whether your strongest passage falls inside the short extract Claude receives. | Rewrite |
| Exact words present | Claude's filter is code, and code matches strings rather than meanings. Whether the distinctive words, product names and figures from your searches are on the page. | Rewrite |
| Short enough to quote | Whether your key sentences fit the short span Claude quotes, and name their own subject. | Rewrite |
| Claude's history with your site | How often Claude has used your site as a source in the answers we track. | Not scored |
| Ready to be quoted | The shared writing check, applied to very short spans. | Rewrite |
How the combined score is formed
Each assistant scores itself first, as a weighted average over its own scored checks: a pass counts 100, a warn 50, a fail 0.
The combined score is a weighted average across the assistants. Weights are equal by default and you can change the mix on the results screen; the score, the range and the verdict recompute together from results already collected.
An assistant that did not finish is dropped from the average, out of the total as well as the top of it. It is never counted as zero, because a run that failed says nothing about the page. There is no floor on the combined score either: a page that does well on three assistants and fails outright on the fourth is reported as a good score with a warning beside it.
Before-and-after comparisons use a content-only sub-score, computed the same way over the rewrite-lever checks alone. It is the only figure that stays comparable between a published page, which has a search ranking, and a draft, which does not.
The range
Every score is shown as a range rather than a single number, because these pipelines do not make the same decision every time they see the same page. The range never narrows past a fixed minimum, and it widens when the judgment-based checks disagree between their two readings.
A move inside the range is labeled no meaningful change, in a neutral tone. It is not colored as an improvement and it cannot produce a ready-to-publish verdict. If a rewrite moved the score by less than the range, nothing has been shown to move.
How recommendations are merged
Four assistants produce four lists. Every finding maps to a shared signal, and the lists merge into one, in three kinds:
- Agreed. Two or more assistants asked for the same thing. These come first, because they pay off everywhere.
- One assistant. Only one asked. Worth doing, weighted by how much that assistant matters to your audience.
- Trade-off. Two asks that cannot both be satisfied in the same span of text.
A trade-off is not resolved by averaging; averaging two incompatible instructions produces a rewrite that satisfies neither. It is resolved by giving each ask a different part of the page. When the exact question and the longer, more specific wording want the same spot, the question goes in the heading and the specifics go in the paragraph under it. The rewrite receives those allocations directly, so the draft honors both sides rather than the last instruction it read.
What a rewrite changes
It will front-load the answer, tighten passages into self-contained units, carry the specifics and exact terms into the part each assistant reads, split sentences that are too long to quote, move the answer inside the read budget, and produce a revised title, meta description and structured data. It comes with a change log tracing every edit back to the finding behind it.
It will not invent facts, statistics, credentials, customers or dates, and it will not add "best", "leading" or similar claims the page cannot support. The rewrite is meaning-preserving. If the page does not contain the number a check asked for, it tells you to add it rather than making one up.
Off-page findings are never rewritten. They appear in the list as context, clearly marked, and sit outside the rewrite's objective. A rewrite cannot change what Brave returns or whether ChatGPT names your brand.
Readiness verdicts
After every score the optimizer says what to do next. The verdict follows the current weights, so changing the mix recomputes it along with the score.
| Verdict | Means |
|---|---|
| Revise again | There is headroom. At least one assistant has a selection check that is not passing. |
| Ready to publish | Every assistant that finished passes its selection checks and the combined score clears the publish bar. |
| Plateaued | The last pass did not move the score beyond its range. Use the best version so far. |
| Held back | The rewrite made at least one assistant meaningfully worse, even though the combined score held up. |
| Pending | Nothing has finished scoring yet. |
A rewrite is supposed to raise the combined score without making any assistant worse, so a version where Claude fell while the average rose is held back and the previous version offered instead. Nothing is lost: the live page is untouched and the draft is still there if you disagree.
What the outline is built from
Scoring answers why a page was not cited. The outline answers what a page should say before it exists.
From a keyword and a page type, the outline resolves the searches to write for, reads what the ranking pages cover and what none of them answers, and fills a page-type template into a section-by-section brief. Each section carries its target searches, a word budget, the terms that have to appear, and the rules that decide whether a passage can be quoted. Score the draft when it is written; it runs against the same searches the outline was built from.
Scoring versions
Every run records the scoring version that produced it, shown next to the score. Scores from different versions are not comparable, and nothing in the product compares them. A run keeps rendering with the checks that produced it, and when a check or threshold changes materially the version increments instead of old runs being relabelled.
Limits
All of this approximates systems we do not have the source code for. The real rerankers are proprietary; we use an open-source model that behaves comparably, so the ordering it produces is the useful output rather than the absolute number. Search results move daily, and Google decides per search whether to show an AI Overview, so an absent overview means the check was skipped, not failed. We fetch pages ourselves as SpyglassesBot rather than as ChatGPT, Googlebot or Claude; we check robots rules for the real crawler names, but a site that treats those crawlers differently will look different to us than it does to them. Extract sizes and snippet lengths measure systems that change without notice, and thresholds and weights are starting values that get recalibrated as real runs accumulate.
Clearing every check does not guarantee a citation. It means the factors we can see and measure are in place. The answers themselves stay nondeterministic.