How we measure AI visibility
Every week we ask the AI engines the questions customers ask, several times each, and record which businesses each answer names, cites, and recommends. This page explains how, what the numbers mean, and where they fall short.
The weekly prompt panel
For each trade and metro we keep a fixed set of questions a homeowner asks: about fifteen to twenty per trade. Some are local ("best roofers in" a named city), some are research and cost questions, and some are comparisons. Each question is written to test one piece of the work: a location page, a price table, a dataset, the business's entity data. When the work ships, the question that tests it shows whether the answers moved.
The questions are localized for each metro from public facts about it, such as its suburbs, utilities, and climate. A question that needs a fact we do not have for a metro is dropped for that metro rather than asked with a blank or another city's detail.
A client's panel can also hold brand questions that name the client. Those are private, and they never count toward share of voice: a question that names a business would inflate its numbers.
We review the panel monthly. As the questions customers ask shift, questions are added and retired, so the report keeps tracking what matters. The questions themselves are not published.
The engines
ChatGPT, Gemini, Perplexity, and Claude are all part of the measurement. We ask each through its maker's API with live web search turned on: OpenAI's Responses API with web search, Anthropic's Messages API with its web search tool, Perplexity's Sonar, and Gemini with Google Search grounding. Where an API accepts a location, we pass the metro's; every question also names its place in words.
An engine is measured once we hold its key and its runs succeed. Our public research, which started in Phoenix, so far reports Claude and Gemini; ChatGPT and Perplexity are being added, and each page says which engines its numbers come from.
Repetitions
AI answers are not deterministic: the same question can get a different answer a minute later. So every question is asked of every engine more than once each week, twice by default, each time as a fresh request with no earlier conversation. We report frequencies across all of those answers, never a single result.
Named, cited, recommended: how answers are read
Each answer is read in two passes against the list of businesses we track in that trade and metro.
Businesses an answer names that we do not track yet are flagged for review, then added to the list, so a new or small business is counted from the week it is added. Every classification keeps the sentence it came from, so a person can check it.
- Named
- The answer mentions the business by name or a known alias, anywhere, in any light. A rules pass finds name matches sentence by sentence; an AI classifier then reads the whole answer and must agree, so a generic word that happens to match a business name is not counted. If the classifier fails, the rules decide alone.
- Cited
- A link shown with the answer points at the business's own website. We reduce each cited URL to its registrable domain and match it to the business's domain.
- Recommended
- The answer presents the business as a pick: in a list of options, as a top choice, or as an explicit suggestion. A passing, negative, or "avoid" mention counts as named, not recommended. The AI classifier decides this.
The Google comparison
Once a month we take a snapshot of what Google shows for the searches a homeowner makes in each trade: the businesses in Google Places' top twenty for any of those searches, and, where the snapshot captures Google's results page, the local pack and the top ten organic results. We compare that list with the businesses the AI engines named in recent weekly runs, in both directions: who Google shows that the engines never name, and who the engines name that Google does not show.
A business is counted only if we were tracking it from the start of the runs compared, and a failed Google lookup is reported with the results. The searches are not published.
Sample sizes and limits
We would rather show a small, honest number than a big, vague one. These are the limits to read every figure with.
- Sample sizes are small
- One trade in one metro is at most a couple of hundred answers a week, and fewer while engines are being added. Every published figure carries its count of answers, questions, or businesses and its dates. Read small bases as signals, not verdicts.
- Failed answers are dropped, and failed runs are not published
- An engine call that fails is retried. An answer that still fails is left out of that week. If more than one in five answers in a run fail, the run is marked failed and nothing from it is published.
- APIs are not the consumer apps
- We ask through each engine's API with search on. The app a customer uses can answer differently, because of their account, location, history, or a different model. The panel measures the pattern, not any one person's screen.
- Engines change
- Models and search behaviour change without notice. A week-over-week move can come from an engine update rather than anything a business did. That is why we look at trends over several weeks.
- Classification can err
- Matching names in free text is hard: shared words, abbreviations, and similarly named businesses. The two-pass extractor and the kept evidence sentences keep errors checkable and rare, not impossible.
- Not every answer cites
- Some answers name businesses without showing a source. We count names and recommendations as well as links for that reason.
What we publish and what we keep private
We publish results: who the engines name in each measured trade and metro, which websites they cite, how often the engines agree, and how Google's list compares, each with its sample size and dates, and as free CSV and JSON datasets under CC BY 4.0.
We do not publish the questions, the Google searches, client brand questions, or the playbook. Our own panel, which tracks how the engines answer business owners' questions about AI visibility, is private too.
The index names every business it measures by its full name, client or not, in the same neutral way. Being a client never changes how a business is counted or shown.
Results describe what was measured, not a promise of ranking. The results are on our research pages and in our datasets. Terms are defined in the glossary.
See what our AI business engine says about your business.
Start with your website. Our AI business engine crawls it the way AI crawlers do and emails you a score out of 100 and the three findings that matter most.
Month to month. Cancel any time. No contract.