PRODUCT · 8 MIN READ

The 28-day loop: shipping a fix and proving it moved

The 28-day loop is: pick one recommendation, ship it on a declared date, open an attribution window scoped to the prompts that change should affect, sample those prompts throughout the window, and read the verdict at day 28 - improved, declined, or no change, where movement under five percentage points is reported as no change. One change per window wherever possible, because two overlapping changes produce one unattributable result. The loop's output isn't a score; it's an answer to did that work.

Most marketing measurement fails at the joint between activity and outcome: work ships continuously, numbers wiggle continuously, and narrative fills the gap. The 28-day loop is our attempt to rebuild that joint for AEO. It's how ClerAEO structures its verdicts, but you can run the discipline with a spreadsheet - the tool automates the loop, it doesn't own the idea.

Step 1: one change, declared out loud

The loop starts when you commit to a specific change with a ship date: rewrote the comparison page's opening to answer the query directly, propagated the canonical entity description to twelve directory listings, published the integration guide. The declaration matters because it defines the hypothesis - this change should affect these prompts - before the data arrives. Attribution decided after the fact always finds a flattering story.

Step 2: scope the prompt set

Every change gets a blast radius. A rewritten comparison page should affect comparison prompts, not your whole footprint. Scoping does two jobs: it makes the verdict sensitive, because diffuse effects vanish when averaged across a hundred unrelated prompts, and it makes the verdict falsifiable, because you said in advance where the effect should appear. If the movement shows up everywhere except the scoped prompts, that's not a win - that's a coincidence wearing your badge.

Step 3: sample through the window

From ship date to day 28, the scoped prompts get run repeatedly against each tracked engine - for us, ChatGPT and Claude live today, more rolling out. Repetition is the whole method: engines are stochastic, and a presence rate estimated from many samples is the only version of the number worth reading. To be explicit about cadence, this is scheduled sampling across the window, not daily refresh - we don't claim that, and for a 28-day judgment you don't need it.

Step 4: the verdict, with a floor under it

At day 28, the window closes and the comparison is made: presence rate on the scoped prompts, before versus after, per engine. Three verdicts are possible - improved, declined, no change - and the no-change band is wide on purpose: under five percentage points of movement is reported as no change, full stop. About half the interesting information is in the non-wins. A no-change verdict on a tactic everyone swears by is worth more than a win, because it tells you where your next window shouldn't go.

What the loop can't do

The loop is slower than shipping constantly and narrating the dashboard. It's also the difference between an AEO program that compounds - each verdict pruning what doesn't work - and one that's a year of activity with a trend line nobody can explain.

See where you stand.

Run a free scan and get your own answer-engine scorecard.