A category of software has appeared over the last two years that tracks how companies show up in AI answers. We do not sell one, and we have no stake in which you choose. We do get asked whether they are worth buying, usually by teams who have already trialled two and cannot tell them apart.

The honest answer is that they measure a real thing accurately, and that the thing they measure is not the thing most teams are actually trying to find out.

What they do well

Nearly all of them work the same way. They run a set of prompts against several AI systems on a schedule, parse the responses, and record whether you were mentioned, whether you were cited as a source, and who appeared alongside you. Then they turn that into a score and a trend line.

That is genuinely useful, and it is tedious enough by hand that automating it is worth money. Specifically, they are good at:

  • Sampling at a volume you cannot match manually. Hundreds of prompts, several systems, every week.
  • Detecting movement. A drop in mention rate is a real signal, and you will see it faster than you would notice it.
  • Competitive share. Who is named, how often, relative to you.
  • Source capture. The better ones record which URLs were cited, which is the most valuable field they collect and the one buyers pay least attention to.

If you have no measurement at all, one of these will tell you more in a week than you currently know.

What the number leaves out

A visibility score is a measurement of one of the three realities in the Digital Reality Gap framework: Internet Reality, what the systems and the market actually find. It is a good measurement of it.

The gap that costs you money is the distance between that and Business Reality, what your company actually is. No tool can measure that side, because no tool knows what you are. It has your website, which is Stated Reality, and Stated Reality is frequently the thing that is wrong.

That produces four blind spots.

A score does not tell you which kind of problem you have. Mention rate of 12% is compatible with an evidence problem, a clarity problem, an intent problem and a structural problem. Those have entirely different fixes and entirely different costs. The number is identical in all four cases.

Prompt sets are a sample, and someone chose it. Whoever wrote the prompts decided what counts as your category. If the market describes your problem in words your team does not use, a prompt set written by your team will miss it, and the tool will report improvement while the actual demand goes elsewhere.

Being mentioned is not being recommended. Appearing in a list of eight is recorded as a mention. So is being named as the expensive option, or the one with the caveat attached. Sentiment features help and are still coarse.

Nothing in the output tells you what to change. This is the substantive one. The tool tells you that you are absent from a question. It cannot tell you that you are absent because the two directories the models trust for that category describe you with a positioning you abandoned in 2023. Someone has to open the sources and read them.

What to establish before you buy

If you are evaluating one, the questions that separate them are not the ones on the comparison pages:

  • Does it record cited source URLs, or only mentions? Sources are where the work happens. A tool that reports mention rate without sources gives you a thermometer and no diagnosis.
  • Can you write your own prompts, in your buyers’ language? Preset category prompts are the fastest way to measure the wrong market.
  • How often does it sample, and does it re-run the same prompt? Answers vary between runs. A platform reporting a single response per prompt per week is reporting noise as trend.
  • Which systems, and how does it handle AI Overviews? Coverage varies a lot, and the differences between systems are usually the most informative part of the data.

Where this leaves the decision

Buy one if you need to watch a number over time, defend a budget, or catch movement early. That is a real job and the software does it.

Do not expect it to answer the question you actually have, which is almost never “how visible are we” and almost always “why are we not the answer, and what would change that”. Answering that means reading the sources the systems cite, comparing what they say against what the company genuinely is, and deciding which of the four kinds of problem you are looking at. That is an afternoon of judgement per finding, and it is the part the category has not automated.

A tool gives you the score. It does not give you the reason, and the reason is what you act on.

Start with what the market can see.

A conversation is enough to know whether there is a gap worth closing.