How we measure AI visibility
Every number on this site comes from the process below. It is published in full so you can judge it, argue with it, or reproduce it. If we can't explain a number, we don't publish it.
1. What we ask
We maintain a fixed set of 489 buying questions — the kind people actually type when choosing software. They follow a small number of patterns (best X software, X vs Y, X alternatives, is X worth it) across 40 categories.
The question set is fixed on purpose. If we changed the questions, scores from different weeks would not be comparable — and a trend nobody can compare is worthless.
2. How we ask it
Each question goes to 2 AI engines through their official APIs, using one prompt that never changes:
No brand names are suggested, no lists are primed, no results are re-ordered. The engine answers exactly as it would for anyone else. Every answer is stored with its date and never edited or deleted — 823 archived so far.
3. How we read the answer
AI answers are structured: numbered lists, bold names, headings. That structure is the recommendation order, so we parse it directly rather than asking another model to interpret it. Extraction is deterministic — the same answer always produces the same result, which means re-running our pipeline can never quietly change history.
Section headings and feature phrases are filtered out. A candidate must be title-cased, four words or fewer, and contain at least one word that isn't generic vocabulary — so Zoho CRM is kept and Advanced Capabilities is not.
Name variants are merged: HubSpot, Hubspot, HubSpot CRM and hubspot.com all resolve to one brand. 5,035 mentions extracted so far.
4. The score
prominence = 1 / (1 + 0.6 × ln(position))
raw = reach × average(prominence)
score = 100 × raw / highest raw in the category
Two things decide visibility: how many questions name you, and how early you're named. Being first in an answer counts roughly twice as much as being tenth (position 1 = 1.00, 2 = 0.71, 3 = 0.60, 5 = 0.51, 10 = 0.42).
The result is scaled so the most visible brand in a category sits at 100 and everyone else is relative to it — the same approach Google Trends uses. It makes the number easy to read and comparable over time.
5. Sample size and honesty
- Scores use a 7-day window, not a single day — LLM answers vary slightly between runs and no engine covers every question daily.
- A brand named in fewer than 4 separate questions gets no public page. The data exists; we just won't publish a score we can't defend.
- Every score is shown with the number of questions behind it. Read those together.
- Engines are also scored separately, because they disagree — and that disagreement is often the most interesting part.
6. What this does not measure
- Not quality. A high score means AI names you often, not that you're the better product.
- Not personalised results. We measure default answers, without accounts, history or location.
- Not every engine. We cover the engines we can reach through official APIs within free limits. Coverage is listed on each brand page.
- Not the whole web. Only the questions in our tracked set.
Found a mistake?
Merged brands, wrong names and bad extractions do happen. Tell us and we'll fix it — corrections make the index more useful, not less credible.