PRODUCT · 7 MIN READ

Why a visibility score without attribution is a vanity metric

A visibility score tells you where you stand; attribution tells you why the number moved. Without attribution, a score that rises can't tell you what to do more of, and a score that falls can't tell you what to fix - every movement is a Rorschach test your team reads optimistically. A metric you can't act on is a vanity metric, however rigorously it's computed. The useful unit isn't the score; it's the change-verdict pair: we shipped this, and here is what it did.

The first generation of AI visibility tools converged on a familiar deliverable: a score, a trend line, a dashboard. We ship a score too, so this isn't a takedown of scores. It's an argument about what a score is for - and what has to sit next to it before it deserves budget.

The failure mode: a number in search of a story

Your score goes from 54 to 61. What happened? Maybe the three blog posts you shipped. Maybe the entity cleanup. Maybe the engine changed its retrieval behavior, or your competitor's pricing controversy pushed them out of answers, or the sampling variance we all live with rolled in your favor. A score alone cannot distinguish these - and teams reliably attribute rises to their own work and falls to the algorithm. That's how a metric becomes a mood.

What attribution means here, concretely

Attribution in AEO can't be causal in the clinical-trial sense - you can't run a randomized experiment on ChatGPT's answers. What you can do is make attribution structural: tie every measurement window to a declared change. In ClerAEO, the working unit is a recommendation you ship on a date, an attribution window that opens on that date, and a measured verdict at day 28 comparing your presence rate on the affected prompts before and after. The score still exists, but movement is always read against a declared change, on a scoped prompt set, over a fixed window.

The verdict has to be allowed to say nothing happened

This is the uncomfortable part, and the part we think matters most. Answer-engine sampling is noisy; small movements are usually noise. So our verdicts have a floor: at day 28, movement under five percentage points is reported as no change. Not modest improvement. No change. This costs us - a tool that regularly tells you your last month of work didn't move the number is harder to love than one that always finds a win. But a verdict system that can't say nothing happened can't be trusted when it says something did.

Honest limits of our approach

The question to ask any tool, including ours

When this number moves, how do I know why? If the answer involves a scoring methodology but no link between changes and windows, you're buying a mood ring with good typography. If the answer includes a mechanism for saying your work didn't move it - that's a tool trying to earn the next decision, not just renew the subscription.

See where you stand.

Run a free scan and get your own answer-engine scorecard.