PLAYBOOK · 8 MIN READ

How to choose an AI visibility tool: an 8-point evaluation checklist

Evaluate an AI visibility tool on eight points: engines live today, sampling depth, who controls the prompt set, what is measured beyond presence, citation provenance, attribution to shipped changes, where recommendations come from, and commercial transparency. Most demos showcase the dashboard, which is the least differentiated part of any of these products. Ask about the measurement underneath it instead.

This checklist is vendor-neutral and includes the questions that are uncomfortable for us. We build an AI visibility tool, and on several of these points a competitor will give you a better answer than we will. That is the point of a checklist - it should sort tools by fit, not by who wrote it.

1. Which engines are live today

Every vendor in this category lists engines. Fewer distinguish between what is queried in production today and what is on the roadmap. Ask for the list of engines currently returning data on a paid account, and ask when each one went live. Then ask what happens to your historical numbers when a new engine is added - if the new engine is folded into a single blended score, your trend line breaks silently at the moment of the upgrade. Demo question: which engines return data on a paid account today, and are per-engine numbers reported separately or blended?

2. How deeply each prompt is sampled

Answer engines are stochastic. The same prompt asked twice can produce different brands, different order, different citations. A visibility number is only an estimate of a rate, and the quality of that estimate depends entirely on how many samples sit behind it. Some tools run each prompt once per period and present the result as a measurement. Demo question: how many times is each prompt run against each engine per reporting period, and is that number the same on every plan tier?

3. Who controls the prompt set

The prompt set is the denominator of everything. If the vendor generates it, changes it silently, or expands it as you upgrade, your numbers stop being comparable to your own history - and a score that rises because the question list got easier is worse than no score. Demo question: can I define and freeze the prompt set, will I be notified when it changes, and can I export the full list with the raw answers behind each number?

4. What gets measured beyond presence

A brand named first and recommended, and a brand named fifth in a but-consider-alternatives clause, are both present. A presence-only tool scores them identically. At minimum you want position (where in the answer you appear, and how often you are the first mention) and framing (recommended, listed neutrally, or used as the unfavourable contrast). Demo question: show me two answers where we appear, one good and one bad, and show me how the product scores each differently.

5. Citation provenance

Some citations are pages the engine actually fetched, verifiable against the live URL. Others are named from training memory and may be stale or wrong. They look identical in the prose. A tool blending them into one number is mixing a population that responds to your page fixes within weeks with one that responds over months or not at all. Demo question: how do you distinguish a retrieved citation from a model-recalled reference, and what do you do with the ones you cannot classify?

6. Attribution: can it connect a change to a movement

This is the point where most tools in the category stop. A score that moves without an explanation cannot direct the next dollar, and teams reliably attribute rises to their own work and falls to the algorithm. Look for a structural link - a change declared with a ship date, a prompt set scoped to that change, a fixed window, and a verdict at the end. Demo question: when this number moves, how does the product tell me why, and can it attribute a movement to a specific change we shipped on a specific date?

7. Where the recommendations come from

AEO advice is full of confident numbers with no source behind them - question-format headings doubling citations, schema types adding thirty percent. We have looked for primary sources for these and cannot find them. A recommendation engine that repeats folklore will fill your quarter with work that cannot pay off. Demo question: pick one recommendation the product would give us and show me the evidence behind it - a study, or a mechanism you can verify on our pages. If neither exists, why is it in the product?

8. Commercial transparency and exit

Roughly half this category publishes no pricing at all. That is not disqualifying, but it changes your evaluation timeline, because every quote becomes a sales cycle. Published entry pricing as of July 2026 gives you anchors - Rankscale from $20/mo, Otterly.AI from $29/mo, Profound from $99/mo, the Semrush AI Toolkit add-on from roughly $99/mo, Scrunch AI from $250/mo annually, AthenaHQ Starter at $295/mo above a free tier, Ahrefs Brand Radar at $398/mo - while Peec AI, Evertune, Brandlight, Bluefish, Writesonic, Daydream and Goodie AI publish none. Pricing and features here are from public sources, verified July 2026, and subject to change. Demo question: what is the total first-year cost including any suite subscription required, what is the contract term, and can I export my full history if I leave?

The one question that sorts the field fastest

Ask: what is your threshold for calling a movement real, and what does the product do when a change did not work? A vendor with a stated threshold has thought about sampling variance and has accepted the commercial cost of telling customers their last month achieved nothing. A vendor without one will show you a chart that always finds a story. For reference, our own answer is that movement under five percentage points at day 28 is reported as no change - which regularly means telling a customer their work did not move the number.

Run this checklist across three vendors and the shortlist usually collapses to one. Most of the differentiation in this category is in the measurement layer, and almost none of it is in the dashboard everyone leads the demo with.

Frequently asked

What should I ask on an AI visibility tool demo?

Which engines return data on a paid account today, how many times each prompt is sampled per period, whether you control and can freeze the prompt set, how position and framing are scored beyond simple presence, how retrieved citations are distinguished from model-recalled ones, how the tool attributes a movement to a shipped change, what evidence sits behind one specific recommendation, and what the total first-year cost and exit terms are.

How much should an AI visibility tool cost?

Published entry pricing as of July 2026 spans $20/mo to $398/mo depending on the shape of the product, and several vendors publish nothing. Budget by the question you need answered: a presence number costs tens of dollars a month, enterprise brand analysis costs a sales conversation, and the total first-year cost of a suite add-on includes the suite subscription itself.

Do I need a dedicated AI visibility tool at all?

Not always. If you already pay for an SEO suite with an AI visibility feature, start there and prove you have outgrown it. Buy a dedicated tool when you need something the suites do not do - per-engine depth, citation provenance, framing analysis, or attribution that ties a shipped change to a measured verdict.

See where you stand.

Run a free scan and get your own answer-engine scorecard.