# Choosing the Right Prompts to Track

Picking the prompts you track is the single biggest decision you'll make in an AI visibility program. Track the wrong ones and you'll measure a market you don't sell into. Track too few and every number will wobble. This guide walks through how to choose a set that actually reflects how your buyers shop, and how to know when you have enough of them.

<Callout>
Prefer a printable version? [Download this guide as a PDF](/downloads/spyglasses-prompt-selection-guide.pdf).
</Callout>

## What You'll Learn

In this guide, you'll learn:

- Why AI visibility is measured like a poll and not like a search ranking
- How the industry is converging on prompt tracking as the unit of measurement
- How to design a prompt panel from your brand snapshot
- How many prompts you need for reliable, directional, or exploratory data
- How to cover buyer segments, problems, and buyer-journey stages
- How locations and personalization change what you should track
- Which prompts belong in a share of voice measurement and which don't
- A summary table you can use as a starting checklist

## Background

Buyers are asking AI before they ask anyone else. It happens on small day-to-day purchases, and it happens on six-figure B2B deals where a committee wants a shortlist before the first call. In both cases the assistant answers with a handful of brands, and the ones it names get the meeting.

That creates a measurement problem. In traditional search you have first-party data. Search Console tells you the exact queries people typed to reach you. Nothing like that exists for AI assistants. The conversations happen inside a private chat, and getting access to them would mean reading people's private conversations, which no responsible platform is going to hand over and no responsible vendor should ask for.

There's a second complication. AI answers are nondeterministic. Ask the same question twice and you can get two different lists of brands. That isn't a bug in the measurement; it's how the models work. So AI visibility measurement behaves much more like audience polling than like rank tracking. You aren't looking up a fixed position. You're sampling a population of possible answers and estimating what share of them mention you.

Once you accept that framing, everything else in this guide follows. A poll gets more accurate as you ask more people. A prompt panel gets more accurate as you track more prompts.

## The Emerging Field of Prompt Tracking

Every organization tracking AI visibility is solving the same equation. More prompts, more platforms, and more locations produce a tighter estimate, and they also cost more to run. So the real question is never "what's the perfect panel"; it's "what's the most accurate panel I can afford".

You can put numbers on that trade-off before you commit. Our [AI visibility data reliability calculator](/is-your-ai-visibility-data-reliable) shows the margin of error you'd get at a given prompt count, so you can see what an extra ten prompts actually buys you.

The industry is settling on this same view. The IAB's [Measuring Visibility in the AI Era](https://www.iab.com/guidelines/measuring-visibility-in-the-ai-era/) guidelines treat AI visibility as a sampled measurement with stated confidence rather than a fixed position. AMEC's [GEO hub](https://amecorg.com/amec-geo-hub/) makes a similar case from the communications measurement side, tying AI visibility back to established standards for evaluating earned media.

Our own research backs it up:

- [Prompt phrasing consistency](https://research.spyglasses.io/articles/prompt-phrasing-consistency) shows that rewording the same question changes which brands appear, so a single phrasing is a single data point and not the answer.
- [Rank stability in ChatGPT](https://research.spyglasses.io/articles/rank-stability-chatgpt) measures how much the brand list moves between identical runs.
- [Synthetic prompt coverage](https://research.spyglasses.io/articles/synthetic-prompt-coverage) looks at how well a generated panel stands in for the real distribution of buyer questions.

Every organization balances accuracy against cost. The rest of this guide is about doing that deliberately instead of by accident.

## How to Design a Prompt Panel

### Start With Your Brand Snapshot

The best starting point is an AI Visibility Report, and at minimum a [Brand Snapshot](/docs/dashboards/brand-snapshot). The snapshot captures the three facts every good prompt is built from:

- **Category**, or what kind of thing you are
- **Target customer segments**, or who buys you
- **Problems solved**, or what goes wrong for someone before they need you

If those three fields are vague, your prompts will be vague. Fix the snapshot first.

### Build Prompts Like Mad Libs

Once the snapshot is right, writing prompts is mostly a fill-in-the-blank exercise. You take a template and drop in the snapshot facts. A "best in category" prompt looks like this:

> What is the best [description of category] for [description of target customer]?

Fill it in and you get something a real buyer would type, like "What is the best appointment scheduling software for independent dental practices?"

Every prompt type has a template like this. Spyglasses generates the first set for you, and you can see the full list of [query types we generate](/docs/dashboards/discovery-queries#query-types-we-generate) along with an example of each. If you want a bigger panel built from established marketing frameworks, the [framework-based generator](/docs/dashboards/discovery-queries#generating-prompts-with-ai-frameworks) will build prompts from Category Entry Points, Jobs to Be Done, the Buyer's Journey, and stakeholder perspectives.

### The Minimum Viable Panel

Start with at least one prompt from each of these five types:

| Prompt type | What it tests |
|---|---|
| **Category Best** | Whether you show up when someone asks for the best option in your category |
| **Use Case Solution** | Whether you show up when someone describes a problem instead of a category |
| **Segment Focused** | Whether you show up for the specific kind of buyer you sell to |
| **Solution Comparison** | Whether you make the shortlist when AI is asked to compare options |
| **Budget Constrained** | Whether you show up when price is the leading concern |

If you don't sell to budget-conscious buyers, swap Budget Constrained for another type that fits your market better. A premium brand is usually better served by a second Segment Focused prompt.

### How Many Prompts Is Enough

Five prompts gives you directional data. It tells you roughly where you stand and it catches large moves. More is better, and the improvement is steep at the low end.

Even a small panel detects drift. You're watching your own share of voice and your competitors' share of voice over time, and when those lines cross or diverge, something changed in the market. That signal shows up long before the absolute numbers get precise.

Spyglasses puts the statistics in front of you rather than making you estimate them. When you assign prompts to a Project, the **Data Accuracy** card shows the margin of error at 90% confidence for your selected set, plus the minimum change you'd be able to detect. The [reliability calculator](/is-your-ai-visibility-data-reliable) does the same math outside the app.

The ratings work like this:

| Rating | Margin of error | What it's good for |
|---|---|---|
| **Reliable** | Under 8% | Reporting numbers to a client or an executive team |
| **Directional** | 8% to 15% | Watching trends and spotting real movement |
| **Exploratory** | Above 15% | Getting a first read on where you stand |

### Use the Coverage Matrix as Your Guide

Once you're past the first five prompts, the question becomes what to add next. The [Coverage Matrix](/docs/dashboards/discovery-queries#measuring-coverage--reducing-redundancy) answers that by scoring your library across three dimensions:

1. **Buyer segment and subsegment.** Every distinct kind of buyer should have at least one prompt written from their point of view.
2. **Problems solved.** Every problem in your snapshot should show up in a prompt that describes the problem without naming a category.
3. **Buyer-journey stage.** Awareness, consideration, and decision, following the standard [buyer's journey framework](https://online.hbs.edu/blog/post/buying-journey).

The matrix shows you the empty cells and prioritizes which to fill first. That's a better use of your next ten prompts than adding more variations of the ones you already have.

### A Note on Awareness-Stage Prompts

Awareness prompts feel wrong the first time you read the answers. A buyer at that stage often doesn't know the market well enough to ask for a comparison. They ask about the problem, not the product, so the answer is an explanation rather than a list of brands.

Don't worry if your brand isn't in those answers. That's expected. What you want to track instead is whether your [key messages](/docs/dashboards/key-messages) appear. If AI explains the problem using the framing your marketing put into the world, your differentiators are already shaping the conversation at the earliest possible stage, which is worth more than a name-check.

Spyglasses handles this for you by separating awareness prompts out of share of voice and reporting them as **Share of Influence** instead. See [SoI vs. SoV](/docs/ai-visibility-features/awareness-tracking#two-different-metrics-soi-vs-sov) for how the two metrics differ.

### Consideration and Decision Prompts

Consideration is where AI builds the shortlist. These are the "best X for Y" and "compare A and B" prompts, and they're where share of voice matters most, because the answer is literally a list of who gets considered.

Decision prompts are narrower. The buyer has mostly made up their mind and is now asking specific questions about your solution and their own circumstances. Does it integrate with the tools they already run? Does it meet their compliance requirements? Can it handle their volume? These prompts often surface a completely different set of competitors than the consideration prompts do, which is exactly why they're worth tracking separately.

### Turn Your Set Into a Project

Once your panel fits both your accuracy target and your budget, set it up as a [Project](/docs/dashboards/projects). Projects are what run your prompts on a schedule with Daily Prompt Tracking, hold your baseline, and give you the trend lines that make the whole exercise useful.

## Tracking Visibility in Multiple Locations

AI assistants personalize by location. Ask for the best accounting software in Berlin and you'll get a different answer than you would in Chicago, even with identical wording.

In Spyglasses you set a location for the brand, either a city or a country, and every platform is then tested with that location as context. If you sell into multiple markets, or you're launching in a new one, you can add more [locations](/docs/dashboards/locations) and track them side by side.

Be deliberate here, because locations multiply. Every prompt in a project runs once per location.

<Callout>
**Watch your run count.** 5 prompts across 4 locations is 20 prompt runs, not 5. A panel that looked affordable in one market can quadruple when you add three more.
</Callout>

If budget is tight, it's usually better to run a smaller panel in each market you actually sell into than a large panel in one market you happen to be headquartered in.

## Tracking Personalization in AI Responses

AI assistants also personalize on memory. If someone has told ChatGPT they run a small nonprofit, later answers lean toward that context, and two people asking the identical question can get meaningfully different lists.

We can't simulate an individual's memory, and for the same privacy reason we can't see real user prompts. Nobody should be building a product on top of other people's private chat history.

The best available approximation is to write prompts aimed at the different customer segments in your brand snapshot. A prompt that says "for a two-person law firm" and a prompt that says "for a 400-attorney firm" produce answers shaped by different assumed contexts, and comparing them tells you which segment you're winning.

## Additional Considerations

### Keep Your Brand Out of Discovery Prompts

Do not put your brand name, product names, or proprietary terms in prompts you intend to use for share of voice. A prompt that names you will almost always return you, and it tells you nothing about whether a buyer who has never heard of you would find you.

The test is simple. A discovery prompt should read like it was typed by a prospect researching a category with no preconception about who's in it.

### When You Do Want to Ask About Your Brand

There's a legitimate reason to ask AI about your brand directly, which is to check how it describes you. Spyglasses has two dedicated prompt types for exactly that:

- **Brand Identity** asks what your brand is, what it does, and whether it's any good.
- **Brand Comparison** puts you head to head with a named competitor.

Both are excluded from share of voice and reported through their own metrics, so using them doesn't distort your discovery numbers. Both also support per-prompt expected messages, so you can list the things you want said about you and the things you don't, and track whether AI complies.

We recommend at least one Brand Identity prompt, plus one Brand Comparison prompt for each key competitor. Full detail is in [Brand Identity and Brand Comparison prompts](/docs/dashboards/discovery-queries#brand-identity-and-brand-comparison-prompts).

## Prompt Tracking Guidance, in Table Form

Use this as a starting checklist. Adjust the core count up as your accuracy target tightens.

| Prompt group | How many | What to use |
|---|---|---|
| **Core prompts** | 5 to 50, depending on the accuracy you need | Any of the Standard AI Visibility types; Local prompts (Local Services, Local Comparison, Local Reviews); Buyer's Journey; Jobs to Be Done; Stakeholder Perspectives; Category Entry Points |
| **Brand Identity prompts** | At least 1 | Add more to test specific features, a launch, a rebrand, or M&A activity |
| **Brand Comparison prompts** | One per key competitor | Pick the competitors you actually lose deals to, not the whole category |
| **Project-specific prompts** | As many as the campaign needs | Prompts tied to a campaign, launch, or PR push. More prompts improve accuracy, but watch for overlapping prompts that don't test anything unique |

On that last point, Spyglasses will tell you when you're duplicating coverage. The [overlap checker](/docs/dashboards/discovery-queries#checking-a-new-prompt-for-overlap) compares a new prompt against everything already in your library before you save it, and redundant-cluster detection flags groups of prompts that are all measuring the same thing. Ten prompts that ask the same question in ten ways cost ten times as much and give you roughly one prompt's worth of information.

## Related

- [Brand Snapshot](/docs/dashboards/brand-snapshot) - The source of every fact your prompts are built from
- [Prompts Dashboard](/docs/dashboards/discovery-queries) - Create, generate, import, and manage your prompts
- [Projects](/docs/dashboards/projects) - Run your panel on a schedule and track it over time
- [Locations](/docs/dashboards/locations) - Track visibility across multiple markets
- [Key Messages](/docs/dashboards/key-messages) - Define the narratives measured at the awareness stage
- [Awareness Tracking](/docs/ai-visibility-features/awareness-tracking) - How Share of Influence works
