PRODUCT · 7 MIN READ
What we get wrong: a methodology note on model-recalled citations
When an answer engine cites a source it actually fetched, that's an observed citation - verifiable against the page. When it names a source from training memory without retrieving it, that's a model-recalled citation: sometimes accurate, sometimes stale, sometimes invented. ClerAEO distinguishes the two where engine metadata allows and flags the rest as unverified, but no tool can fully separate them from outside the engine. This note documents where our measurement is solid, where it's inference, and what we refuse to claim.
Every measurement product has a section of the methodology it would rather not write. This is ours. We think publishing it matters more in AEO than in most categories, because the discipline is young, the vendors are loud, and the failure mode - confident numbers built on unexamined assumptions - is the same one we criticize in others.
Two kinds of citation, one kind of answer
When ChatGPT or Claude browses, retrieved URLs are part of the answer's visible apparatus: the engine fetched a page and points at it. Those observed citations are the solid ground - we can check that the page exists, mentions you, and says roughly what the answer attributes to it. But engines also produce answers without retrieval, from training memory, and those answers name sources too: a brand, a review site, according to industry analyses. These model-recalled references look identical to real citations in the prose. They may reflect a page that existed years ago, a page that says something different now, or nothing that ever existed.
Why this matters for every visibility number you've seen
A tool counting mentions and citations across answers is silently mixing two populations: references anchored to live retrieval, which respond to page changes within weeks, and references anchored to training data, which respond slowly or not at all. Blend them and you get numbers that mislead in both directions - content fixes look weaker than they are, diluted by recall-based answers they can't touch, and stale recalled mentions of your brand prop up a presence rate that no current work is earning. Any vendor showing you one blended number has this problem, whether or not they mention it.
What we actually do, and where it's inference
- Where engine responses expose retrieval metadata, we separate observed citations from recalled references and verify observed ones against the live page. This is the reliable part.
- Where metadata is absent or ambiguous, we classify from answer signals - browsing indicators, citation structure. This is inference, it has an error rate, and those references are flagged unverified rather than folded into verified counts.
- Verdicts on content recommendations are weighted toward retrieval-backed prompts, since that's the loop a page fix can plausibly move within a 28-day window.
- This holds for the engines we track live - ChatGPT and Claude - with Perplexity, Gemini, Google AI Overviews, Google AI Mode and Copilot rolling out. Each engine exposes different metadata, so classification confidence differs by engine, and we'd rather say so than average it away.
What we refuse to claim
We can't tell you why a model recalled what it recalled - training data isn't inspectable from outside. We can't promise that fixing your pages fixes recalled answers on any schedule; training-side lag is real and we don't control it. We don't track Google AI Overviews, we don't refresh daily, and we don't publish lift multiples, because each of those claims would outrun what our measurement can support. When a movement is under five points at day 28, we say no change even when the story writes itself.
Why publish this
Because the alternative is the industry default: precision theater, where every number arrives without confidence intervals or caveats and every case study finds a win. Our bet is that operators would rather have a smaller number they can defend in a budget meeting than a bigger one they can't. If reading our limits made you trust the rest of our numbers less, the methodology is working - that skepticism is the correct posture toward every AEO number, including ours.
See where you stand.
Run a free scan and get your own answer-engine scorecard.