RESEARCH · 9 MIN READ
How answer engines choose sources - what the research actually shows
The supportable picture: answer engines lean heavily on conventional search to pick sources, cite only a fraction of what they retrieve, and reward pages with extractable, evidence-backed claims. The GEO paper showed content changes can lift generative visibility by up to 40%; Ahrefs' 1.4M-prompt study showed search-sourced content dominates citations. Beyond that, most specific numbers you'll read - heading tricks, schema multipliers - have no primary source we can find.
AEO is young enough that a handful of real studies carry the whole evidence base, and long enough in the tooth that folklore has calcified around them. This post separates the two. We cite what we can, and we name what we can't.
Finding 1: content-side changes measurably move visibility
The foundational work is the GEO paper (arXiv:2311.09735) from researchers at Princeton, Georgia Tech, IIT Delhi, and AI2, published at KDD 2024. They tested nine content modifications against a benchmark of generative-engine queries and found the best-performing changes - adding citations, quotations from sources, and statistics - improved a page's visibility in generated answers by up to 40%. Two things practitioners routinely drop from that sentence: up to is a ceiling, not an average, and the result comes from a benchmark, not a guarantee about your pages. Still, the direction is solid: evidence-dense pages get surfaced more than stylistically identical pages without evidence.
Finding 2: retrieval runs through search, and citation is selective
The largest observational dataset is Ahrefs' study of 1.4 million prompts. Three results matter. First, ChatGPT cites roughly half of the URLs it retrieves - so getting fetched is the entry ticket, not the prize. Second, search-sourced content is cited 88.5% of the time, against 1.9% for Reddit - the citation layer is, in practice, downstream of conventional search rankings. Third, pages with natural-language URL slugs were cited 89.8% of the time versus 81.1% for non-natural slugs - a real gap, though small enough that nobody should rebuild their URL structure expecting a transformation.
Three popular claims we cannot source
- Question-format headings get 2x citations. Repeated across dozens of AEO articles; none links a study. We can't find a primary source.
- Consistent heading levels produce 40% more citations. The number appears to be a mutation of the GEO paper's unrelated 40% figure. No study we can find tests heading levels at all.
- FAQPage schema adds 30% more citations. Schema markup helps search engines parse pages, but we can find no controlled test tying any schema type to a citation lift of any size.
None of these practices is harmful - clear headings and structured markup are fine hygiene. The problem is the precision. A made-up multiplier gets budget allocated and expectations set, and when the lift doesn't arrive, the whole discipline loses credibility. If someone quotes a number, ask for the study. That one habit filters most of what's wrong with AEO content.
What the research doesn't cover
Both studies examine citation behavior - answers with retrieval in the loop. Neither tells you much about model-recalled answers, where the engine answers from training memory with no sources fetched. Those answers move on slower timescales and respond to different work, mostly entity presence and consistency across the corpus. Public research here is thin, and we'd rather flag the gap than fill it with confident guesses.
What we do with this at ClerAEO
Our recommendations engine only ships advice in two categories: things with a primary source (the two studies above, and their successors as they appear) and things we can verify mechanically on your pages, like whether an answer exists above the fold or whether your entity descriptions agree. If a tactic has neither a study nor a mechanism, it doesn't go in the product. That rules out some popular advice, and we're comfortable with that.
See where you stand.
Run a free scan and get your own answer-engine scorecard.