An AI visibility tool runs a fixed set of prompts against AI answer engines on a schedule, records whether your brand was named or cited in the generated answer, and keeps the history so you can read a trend instead of a snapshot. This guide is for the growth lead or in-house SEO who has already run the manual check, watched the brand appear in some answers and not others, and now has to justify a subscription to someone who will ask what the number on the dashboard actually means. It covers where these platforms get their data, why two of them can report different figures for the same brand in the same week, why those figures are rarely comparable even when they look alike, and the questions worth asking before you sign anything. We do not sell a tool in this category, which is worth stating plainly: when we pulled the organic results for this query in August 2026, five of the nine were published by companies selling a product in the category they were ranking for.
The short version
- The percentages these tools report are not comparable with each other. Each vendor defines its denominator differently, tracks a different engine set on the plan you will actually buy, and publishes a different amount about how it samples.
- The category has converged on the same four metric names: coverage, share of voice, citation quality and sentiment. The names match across vendors; the formulas underneath them do not.
- Two tools can report different share-of-voice figures for the same brand in the same week and both be correct. Non-determinism, per-engine difference and personalization are three separate causes, before the definitional gap above is even counted.
- Buy when manual tracking breaks on one of four specific axes, and not before. Our guide to how to track AI search visibility sets out those axes; this guide starts after you have hit one.
What is an AI visibility tool, and what does it actually do?
It automates a question you could ask by hand. The tool holds a list of prompts, sends them to one or more answer engines on a schedule, parses each generated answer for your brand name and your URLs, and writes the result to a history you can chart. Everything else on a vendor's feature list, the sentiment scoring, the competitor benchmarks, the prompt research, is built on top of that single loop.
Two adjacent categories get confused with it, and the confusion is expensive because all three are sold to the same buyer.
| Tool type | What it records | Where it stops |
|---|---|---|
| Rank tracker | The position a URL holds for a query on a results page | It cannot see inside a generated answer, where there are no positions to take. |
| Brand monitoring | Mentions of your name across the open web, news, forums and social | It watches the inputs that shape AI answers, not the answers themselves. |
| AI visibility tool | Whether a generated answer named or cited you, for prompts you chose | It sees the answer, not the person who read it. |
The metric definitions sit upstream of the purchase decision. Coverage rate, share of voice, citation quality and sentiment are the four readings this category reports, and our guide to what to measure in AI search defines each one and argues the order to trust them in. What follows is about the instrument, not the reading.
How do these tools get their data?
By running prompts and keeping the answers. The interesting question is how they run them, because there are two broad routes and the choice changes what you are buying more than any feature comparison will.
Official API access: the vendor calls a model provider's API and stores the response. This is stable, contractually permitted, and it scales. It also does not always return what a person sees. An API call to a model is not the same product as the consumer surface built on top of it, and Google publishes no API that returns AI Overview text, so any tool reporting on that surface is reaching it some other way.
Reading a consumer interface: the vendor drives the product a real user would use and captures what comes back. This is closer to what your buyers experience and it is considerably more fragile. It depends on a surface the vendor does not control, and it can be rate limited, redesigned or blocked without notice.
Most platforms use some mix of the two, and none of the three pages we read sets out, engine by engine, which route it uses. Access terms in this category are genuinely unsettled, so treat the gap as a question to ask rather than something to resent: how, specifically, do you reach each engine you list?
Where a vendor does describe the pipeline, the description is useful. Ahrefs states publicly that Brand Radar starts from real queries in its own keyword database, expands them into natural questions, and runs the resulting prompt set through each AI platform, storing the responses. That is a description of a method rather than a description of an outcome, and it is the kind of statement worth looking for before you compare two dashboards.
Why do two tools report different numbers for the same brand?
Three causes, and they stack. In our AEO and GEO work, the same brand can look present and absent inside a single afternoon, so a disagreement between two dashboards is usually not a bug in either of them.
- Non-determinism: the same prompt returns different wording, different sources and a different cast of brands on two consecutive runs. Every run is a sample, not a reading, and two tools sampling independently will land in slightly different places even when nothing has changed.
- Per-engine difference: ChatGPT, Perplexity, Google AI Overviews, Google AI Mode, Gemini, Claude and Copilot build answers differently and cite differently. A tool that weights AI Overviews heavily and a tool that weights ChatGPT heavily are describing two different worlds, both accurately.
- Personalization: the third cause, and the one least often named. Answer engines adapt to session history, account state and location, so the answer a tool receives and the answer you get when you spot-check it on your own laptop can both be genuine and still disagree. Otterly.ai is unusually direct about this in its own materials, telling users that a manual search on ChatGPT or Perplexity may return a different result from what the platform reports, because those platforms store information about you and your search history and personalize on signals such as location.
The consequence is the rule that should govern every figure you are shown in a demo. A share-of-voice number is only comparable to another number from the same denominator: same prompt set, same competitor list, same engine mix, same run count. Change any one of those and you have two different measurements wearing identical notation.
The published definitions make the point better than an argument does. Ahrefs defines AI Share of Voice as the percentage of AI responses within your topic set that mention or cite your brand versus competitors. Otterly.ai describes its Share of AI Voice as the percentage of citations you own versus competitors. Profound's documentation splits the idea in two, calculating a visibility score as the responses including your brand divided by the responses including at least one brand, and share of voice as the responses that mention your brand divided by all brand mentions across responses.
Read those three carefully and you have four denominators: responses in a topic set, citations, responses containing at least one brand, and total brand mentions. None is wrong. Most are called some version of share of voice, and no two belong on the same chart. That, rather than feature depth, is what makes cross-vendor comparison here so slippery.
How many runs does it take before the number means anything?
Enough that the vendor should be able to tell you, and the useful discovery is that at least one of them already has. From my experience, buyers in this category ask about engine coverage inside the first five minutes of a demo and almost never ask how many times each prompt was run. That is the wrong order. Engine coverage tells you what the tool looked at. Run count tells you whether what it saw was a measurement.
The reason it matters: because an answer engine is non-deterministic, every run is a sample rather than a fixed reading. A percentage built from a handful of runs per prompt and a percentage built from thousands are different objects presented in identical notation, and the smaller the sample, the more of the movement on your chart is the instrument rather than the market.
The category handles that unevenly, and Profound disclosed the most of the three when we checked in August 2026. Its documentation states that tracked prompts run daily against each configured platform, its pricing lists a monthly response count per plan, and it has published an analysis of its own sampling frequency arguing that, for visibility, once a day already lands "within about 2 percentage points" of a ten-times-a-day reading. We have not audited that analysis. What matters for a buyer is that it exists in public, with its reasoning attached.
You still need to know what a sample can and cannot support, because the plan you buy may be a great deal smaller than the study that defends it.
- A small sample supports presence and absence: if you never appear across a full run of the prompt set, that is a real finding. Zero is a reliable result.
- A small sample supports direction across repeated runs: four consecutive weeks moving the same way is a signal, even when each individual week is rough.
- A small sample does not support a decimal place: a coverage rate quoted to one decimal implies a precision the underlying sampling almost certainly does not have.
- A small sample does not support a small delta: a two-point month-on-month move is the first thing a dashboard will show you and close to the last thing you should act on.
- A small sample does not support a close competitive gap: "we are three points behind them" is a coin flip until you know the run count behind both figures.
None of this argues against buying. It argues for reading the number at the resolution it has earned. So the question to put to a vendor is not whether they disclose their sampling, because some plainly do. It is narrower: on the plan I am buying, how many responses sit behind one reported percentage, on one engine, in one period?
The answer changes by tier, and so does whether you can work it out at all. Where a vendor publishes a prompt count, an engine count, a response total and a cadence on the same row, the per-engine figure falls out of the arithmetic: on a plan tracking 100 prompts against three engines at a daily cadence, a 9,000-response monthly total is one run per prompt per engine per day. Where a vendor publishes a cadence but no response volume, or a response volume but no cadence, it does not, and the pricing page will give you an estimate at best. That is the question to ask, and it is worth knowing before the call whether the answer is already public.
What do the tools in this category converge on?
The vocabulary, and less than you would hope beyond it. Coverage, share of voice, citation quality and sentiment appear across the whole category, which is why the framework is worth learning once with a spreadsheet before you pay anyone to automate it. What does not converge is the arithmetic underneath. As the published definitions above show, the same four words can sit on top of four different denominators.
What separates the platforms sits underneath the metrics: which engines the plan you actually buy reaches, how often it samples them, and how each vendor defines the denominator of its headline percentage. The table below is a snapshot of publicly stated capabilities, read from each vendor's own marketing, pricing and documentation pages on 29 August 2026. It is not a ranking, it is not a test, and we have not trialled these products. Pages in this category are rewritten often, so check the current wording before you rely on any row.
| Platform (as stated, August 2026) | Engines named, and what a plan actually tracks | What the vendor publishes about sampling and metrics |
|---|---|---|
| Profound | Marketing pages name ChatGPT, Perplexity, Claude, Gemini, Grok, Microsoft Copilot, DeepSeek and Google AI Overviews. Pricing is narrower: ChatGPT alone on Starter, three engines on Growth, up to nine on Enterprise. | The most disclosed of the three. Prompts run daily against each configured platform, pricing publishes a monthly response count per plan (1,500 on Starter, 9,000 on Growth), the knowledge base gives the visibility and share-of-voice formulas, and a public analysis defends the daily sampling rate. |
| Otterly.ai | Marketing pages name ChatGPT, Google AI Overviews, Google AI Mode, Gemini, Perplexity, Copilot and Claude. Every plan tracks the same four, with Claude, AI Mode and Gemini sold as add-ons. | Partly disclosed. Daily tracking frequency and a prompt count on every plan, a stated Share of AI Voice definition (citations you own versus competitors), and an unusually candid note on why its results differ from a manual search. No response total published, so the same arithmetic yields an estimate rather than a confirmed figure. |
| Ahrefs Brand Radar | Google AI Overviews and AI Mode, ChatGPT, Perplexity, Gemini, Copilot and Claude. Tracked AI prompts per plan are small, in the 5 to 20 range, with custom prompt tracking sold as a separate add-on. | The most disclosed on provenance. It states where the prompts come from, publishes a total monthly prompt volume for its research corpus, and states that question sets are re-tested monthly on a 90-day reporting window. No per-prompt response count. |
Two things stand out, and neither is what we expected before checking. The first is in the second column: on two of the three, the engine list on the marketing page is a superset of what the plan you are likely to buy actually tracked when we checked. That is not deception, it is how tiered pricing works. But the logo grid driving most of the comparison content on this topic is describing an enterprise contract rather than the product you will use.
The second is that disclosure here is real and wildly uneven, which is harder on a buyer than blanket secrecy would be. One vendor publishes a variance analysis, another a cadence and prompt counts but no response total, a third corpus provenance and a re-test window. There is no shared unit, so the three cannot be lined up and compared without normalising first, and normalising is work the guides ranking for this query did not do when we pulled them in August 2026.
One limit, stated plainly: the table reflects public marketing, pricing and documentation pages on 29 August 2026, not a support conversation or a contract. We did not test the products and did not verify any vendor's statistical claims about its own sampling. It records what each company puts in public, which is the only part a buyer can check before a call.
What should you ask a vendor before you sign?
Seven questions, and the useful part is knowing what a weak answer sounds like. Send them in writing before the demo, because a demo is built to answer a different set of questions entirely.
- Engine access: which engines do you cover, and do you reach each one through an official API or by reading the consumer interface? A weak answer is a logo grid with no method behind it.
- Runs per prompt: how many times is each prompt run in a reporting period, and is that consistent across engines? A weak answer is "continuously", or any reply that avoids a number.
- Raw answers or scores: do you store the full generated answers, or only the derived scores? A weak answer is scores only. You cannot audit a percentage you are not allowed to open.
- Export on exit: if we cancel, can we take the prompt set and its history with us, in a usable format? A weak answer is a PDF report.
- Share of voice, defined: in your own words, what is the numerator and what is the denominator? A weak answer is one that cannot name its own denominator, or that assumes there is only one.
- Sentiment method: is sentiment model-scored or human-reviewed, and against what rubric? A weak answer is "AI-powered". Sentiment is the metric most likely to be quoted in a board deck and the one most easily swung by a single cited source.
- Refresh cadence: how often does the data update, and does the cadence differ by engine? A weak answer is a single number with no per-engine detail behind it, since access routes differ and a uniform cadence is a claim worth probing.
You will not get a clean answer to all seven, and should not expect one. Three or four straight answers plus an honest "we do not publish that" beats seven confident non-answers. What you are testing is whether the vendor thinks in measurements or in dashboards.
When is a spreadsheet still the right answer?
More often than the category would like. A tool buys automation and stored history; it does not buy insight you could not otherwise reach. The four conditions that make manual tracking stop working are set out in our guide to when a paid tracking tool earns its price, and the test is unsentimental: if none of them applies to you yet, a tool is a convenience rather than a need.
There is a second reason to run the manual version first, and it is the one this guide rests on. Building your own prompt set teaches you what a denominator is. Once you have chosen the questions, decided who counts as a competitor, and run the set enough times to watch it wobble, no vendor's percentage will look like a fact to you again. That scepticism is the most valuable thing the free version buys.
What can none of these tools see?
The person who read the answer, joined up to the answer they read. Every platform in this category measures the output side, meaning whether a generated answer named or cited you. The channel itself is not invisible: chatbot referrals carry a referrer, ordinary analytics will show them, and Ahrefs reports AI chatbot traffic in its Web Analytics product, though not in Brand Radar. What none of them joins is a specific mention to a specific arrival, and on Google's own AI surfaces even the aggregate is not separable.
Google's own documentation is direct about it. Sites appearing in AI features are included in overall search traffic in Search Console and reported within the Web search type rather than separated into their own. The same page is equally direct that there is nothing special to do for these surfaces: "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary."
Two things follow. First, the work that earns AI citations is the work that earns rankings, which is why getting cited by ChatGPT, Gemini and Perplexity is not a separate discipline with its own rulebook. When coverage is the problem, the fix is usually upstream of any tool: our guide to diagnosing a citation gap covers the page side, and why AI answer engines name some brands and not others covers the input side that no tracking subscription will fix for you.
Second, be careful what you promote this reading to. We have argued elsewhere that AI-search visibility is not yet a KPI, and nothing here changes that. Choosing a good instrument is separate from deciding the reading should carry accountability. Measure it, learn from it, and keep the number your team is judged on closer to the outcome you want. If the gap is on the execution side rather than the measurement side, that is what our AI search optimization work covers.
The habit worth building is a small reordering of how you evaluate anything in this category: settle the denominator before the demo, not after it. Engine logos are easy to collect and easy to compare, which is why every guide compares them. How a percentage is calculated, across which engines, from how many runs, is what decides whether the chart you are about to build a quarterly narrative on is describing your market or describing its own noise.
This week, do one thing. Open your current tool, or the platform at the top of your shortlist, and find two facts: the formula behind its headline percentage, and how many engines your actual plan tracks. Look on the pricing page and in the documentation rather than the homepage, because the homepage usually carries the full engine list and the pricing page carries yours. If ten minutes does not surface both, that gap is your first question on the next vendor call.
Need a Content Strategy That Actually Ranks?
We help businesses build topical authority with SEO-driven content that performs in both Google and AI search.
Let's Build Your Content StrategyFrequently Asked Questions
What is an AI visibility tool?
An AI visibility tool runs a fixed set of prompts against AI answer engines such as ChatGPT, Perplexity, Gemini and Google AI Overviews on a schedule, records whether your brand was named or cited in each generated answer, and stores the history so you can read a trend. It differs from a rank tracker in what it measures: inclusion inside a generated answer rather than a position on a results page.
Why do two AI visibility tools show different numbers for the same brand?
Three causes, and they stack. AI answers are non-deterministic, so every run is a sample rather than a fixed reading. Each engine builds and cites answers differently, so a tool weighted toward AI Overviews and one weighted toward ChatGPT are describing different worlds. And answer engines personalize on session history, account state and location, so a tool result and your own spot-check can both be genuine and still disagree.
How is AI visibility tracking different from brand monitoring?
Brand monitoring watches the inputs: mentions of your name across the open web, news, forums and social. AI visibility tracking watches the output, meaning whether an AI answer engine actually named or cited you in a generated response to a prompt you chose. The two are related, because the mentions brand monitoring finds are part of what shapes those answers, but they measure opposite ends of the same chain.
Why do traditional SEO tools miss AI search visibility?
Because they measure a different unit. A rank tracker reports the position a URL holds for a query on a results page, which is relatively stable and roughly the same for everyone. An AI answer is generated fresh each time and has no positions to take, so presence, citation and framing replace rank. Measuring it needs a prompt set run repeatedly, not a single lookup.
How many runs does an AI visibility tool need before its number means anything?
There is no single threshold, and disclosure varies more than buyers expect. Because answers are non-deterministic, each run is a sample. A small sample can support presence, absence, and direction across repeated runs. It cannot support a figure quoted to one decimal place, a two-point month-on-month change, or a close competitive gap. Ask what the run count is on the plan you are buying, not the plan in the case study.

Anshuman Sinha
AI SEO Specialist, GrowthHasten
Anshuman Sinha is an AI SEO Specialist and Computer Science Engineer with over three years of experience in SEO and five years in web development. He specializes in Technical SEO, AI Search Optimization (AEO and GEO), SaaS SEO, and building high-performance websites with modern technologies.
View profile



