You Can Track Claude Through the API. Just Not With Sonnet.

Jim Wrubel
10/2/2026
Claude is the easiest AI assistant to leave out of visibility tracking. There's no logged-out version of claude.ai to sample, so every tracked answer is a paid API call, and that cost grows with every prompt, every location and every day. When Claude does get tracked, the obvious way to keep the cost down is a smaller, cheaper model.
That leaves a growing blind spot. Claude's share of worldwide visits to generative AI chatbot websites grew from about 2% to close to 9% between June 2025 and May 2026, according to Similarweb. Its users also look a lot like the people B2B brands sell to. Among the roughly 9,700 Claude users in a recent Anthropic survey, about 30% worked in computer and mathematical jobs, against 4% of US employment, and 23% worked in management, against 7%.
So our research team tested the two questions that decide whether Claude tracking through the API can be trusted. Can an API call stand in for a real person using claude.ai? And do the cheaper models work as well? The answer to the first is yes, if you use Opus 5.5, the model claude.ai now uses by default. The answer to the second is no. Sonnet and Haiku named a measurably different set of brands. The full study is pre-registered and the data is public. Here's what it means for you.
Key Takeaways
- Claude can be tracked through the API, as long as the call uses Opus 5.5, the model claude.ai uses by default. Opus 5.5 came close to claude.ai on which brands it named.
- Sonnet and Haiku are not stand-ins for claude.ai. Every Sonnet and Haiku setup we tested named a different set of brands than claude.ai did for the same B2B software question.
- The cheaper models barely save money. Our estimates put Opus 5.5 with no system prompt at about the same cost per call as Sonnet.
- Adding a leaked copy of claude.ai's system prompt made Opus 5.5 much more expensive without a meaningful change in which brands it named.
- The reasoning setting in claude.ai doesn't change the brands. The higher setting searched more often but named the same brands.
- Spyglasses now tracks Claude on Opus 5.5, using the setup this study found closest to claude.ai.
The study
Every Claude visibility number rests on a choice: which model, with which settings, stands in for a real person typing into claude.ai? We wanted to measure that choice, with costs.
The team wrote 40 buying questions across 20 software categories, such as CRM, payroll and endpoint security. Half ask for a shortlist for a described company. The other half ask how the top options compare for a described need. One person asked every question by hand in claude.ai on three days (September 27, September 29 and October 1), from a fresh Pro account with memory and personalization off.
On the same days, the same questions went through seven API setups: Opus 5.5 and Sonnet 5, each with no system prompt and with a leaked copy of claude.ai's system prompt from a public collection; a low-effort Sonnet variant; Haiku 4.5; and our own production request. In total, 1,080 answers were evaluated in this study.
We compared which brands each answer named and which websites it cited, using an overlap score. Think of it as the share of brands two answers have in common, from 0 (nothing shared) to 1 (the same brands).
The bar was claude.ai itself. claude.ai doesn't give the same answer twice, so the first step was measuring how well it agrees with itself on the same question from one day to the next. Its brand overlap with itself was 0.643. An API setup passes if it agrees with claude.ai about as well as claude.ai agrees with itself. We call the shortfall the gap. Before collecting any data, we set the tolerance at 0.10. For two answers that together name about ten brands, that's roughly one brand swapped for another.
What we found

| Setup | Brand gap to claude.ai | Result |
|---|---|---|
| Opus 5.5, leaked claude.ai prompt | 0.042 | Small, inside the tolerance |
| Opus 5.5, no system prompt | 0.067 | Small, inside the tolerance |
| Sonnet 5, leaked claude.ai prompt | 0.097 | Real gap, at the edge |
| Sonnet 5, no system prompt | 0.119 | Real gap |
| Sonnet 5, leaked prompt, low effort | 0.128 | Real gap |
| Sonnet 5, our former production request | 0.165 | Real gap |
| Haiku 4.5, leaked claude.ai prompt | 0.224 | Real gap |
Sonnet and Haiku name different brands
Every Sonnet 5 setup named a measurably different set of brands than claude.ai did on the same question. That held with no system prompt, with the leaked claude.ai prompt, at low effort, and for our own production request. Haiku 4.5 was further off still.
The closest Sonnet setup needs care. With the leaked prompt, its gap was 0.097, right at the edge of the tolerance, and the data allow a gap as large as 0.15. That's a failure to match, not proof of a large gap. But a stand-in has to show that it's close enough, and this one couldn't.
The model itself matters. When Sonnet 5 and Opus 5.5 got the identical request with no system prompt, Sonnet's brand overlap with claude.ai was 0.052 lower. The model alone accounts for a measurable part of the gap.
Opus 5.5 through the API comes close
Both Opus setups landed inside the tolerance, at 0.067 with no system prompt and 0.042 with the leaked prompt. The difference between those two was 0.025, small and inside the tolerance. The leaked prompt mostly adds cost.
It did help with one thing: the order of the brands. On a score that gives more weight to the brands named first, Opus with no system prompt had a gap of 0.097, a real gap right at the edge. With the leaked prompt, the gap was 0.044.
That makes Opus 5.5 through the API a reasonable substitute for claude.ai. It comes close on which brands are named, and a little less close on their order.
The reasoning setting doesn't change the brands
claude.ai lets subscribers raise reasoning from Medium, the default, to High. High searched about twice as often, 2.5 searches per answer against 1.2, and cited 6.8 websites per answer against 4.4. It named the same brands. Our pre-registered test found the difference between the two settings small enough to treat as none. To match claude.ai's brands, a tracker doesn't need to match its reasoning setting. It does need to match the model.
The cost surprise

The usual reason to use a smaller model as a stand-in is cost. In this study, the saving was small. These are estimates, computed from token and search counts at Anthropic's list prices for the Batches API, not taken from a bill.
| Setup | Estimated cost per call (USD) |
|---|---|
| Sonnet 5, no system prompt | 0.065 |
| Opus 5.5, no system prompt | 0.071 |
| Sonnet 5, our former production request | 0.081 |
| Opus 5.5, leaked claude.ai prompt | 0.128 |
The setup that came closest to claude.ai without the leaked prompt was barely more expensive than Sonnet, and it was cheaper than the request we used to run.
What this means for your Claude tracking
Ask which model your Claude numbers come from. If it isn't Opus 5.5, are the brands in your report the brands a claude.ai user sees by default? In this study, no Sonnet or Haiku setup we tried got there.
Read cited sources across many runs, never one. Sources move a lot, even inside claude.ai. For the same question on two different days, claude.ai's cited websites overlapped only 0.145 on the same 0 to 1 scale. Opus 5.5 with no system prompt overlapped claude.ai's same-day sources at 0.126, and the Sonnet and Haiku setups at 0.069 to 0.090. Treat the source results as directional. Any single run tells you little about which sources Claude cites.
Watch the order, not just the list. A brand named first reads differently from a brand named sixth. Opus 5.5 with no system prompt comes close on the set of brands and slightly less close on the order. On order, the Sonnet setups had gaps of 0.136 to 0.176.
What we changed
Spyglasses now tracks Claude on the setup the study validated: Opus 5.5 (claude-opus-5-5) through the API, at medium effort, with no system prompt. The tracked question goes to Claude exactly as written. Claude can run up to 10 web searches per answer, and the location you track is passed only to Claude's search tool. Our former request ran on Sonnet 5 with our own system prompt, and it is the request the study measured at a brand gap of 0.165.
What this doesn't mean
The claude.ai side is one fresh account in one city, Pittsburgh, with memory and personalization off, so the results describe claude.ai without personalization. The study covered 40 synthetic B2B software buying questions over three days, and the results may not hold for consumer products, local services, other languages or longer periods.
The full method, the pre-registered plan, the cost estimates and the released data are all in the research article.
Claude is included in every AI Visibility Report, and Claude Daily Prompt Tracking is available as an add-on on request. Both now run on Opus 5.5. Our data collection methodology explains how we collect from each AI platform. To add Claude to your Daily Prompt Tracking, contact us.