A brand can be cited by an AI assistant this week and absent next week without a single byte on its site changing. On a 10 prompt sample, 1 answer flipping moves the headline share by 10%. Model updates, index refreshes, and the sheer variability of generated answers all move citation share on their own. That instability is the central measurement problem in answer engine optimization, and most reporting in the category quietly ignores it.
The consequence is a category where almost any engagement can be made to look successful in hindsight. Run enough prompts, measure enough weeks, and something will have gone up.
Retrospective measurement is not measurement
The standard report shows citation share across a prompt set, before and after. It looks rigorous. It usually is not, for one reason: the prompt set and the success metric were both chosen after the work, when the results were already visible.
That is not dishonesty in most cases. It is the natural result of nobody being asked to commit in advance. Given 40 prompts and 12 weeks, a genuinely ineffective programme will still produce a chart with an upward line somewhere in it. With 40 prompts sampled weekly that is 480 measurements to choose a story from.
The fix is unglamorous and it is the whole difference. Declare the prompt set, the baseline, and the target lift before any work starts. Citon states its outcome unit that way, as a citation-share lift declared in advance, which is a materially harder commitment than reporting what moved.
A buyer can enforce this without any technical knowledge. Ask for the prompt list and the target number in writing at the start. A vendor that will not commit is telling you what its reporting is going to be worth.
Baselines have to account for natural variance
Declaring a target is not enough if the baseline is a single measurement.
Citation share for the same prompt varies run to run. Measuring once, doing work, and measuring once more produces a difference that includes an unknown amount of noise, and on a small prompt set the noise can exceed the effect entirely.
A usable baseline is repeated: the same prompts, sampled several times across a period before work begins, so the natural range is known. Then a post-work measurement can be read against that range rather than against a single point. A lift inside the pre-existing variance is not a result, however much it cost.
This is also why very small prompt sets are a warning sign. 10 prompts produce noisy percentages that swing wildly on 1 changed answer. The number needs to be large enough that a single flip does not move the headline.
Monitoring and attribution are different purchases
Watching citation share is cheap and getting cheaper. Entry-level monitoring starts around $99 a month, and any competent AEO engagement should include it rather than bill for it. Against a $6,000 retainer floor that is under 2% of the spend, which is roughly the share of it that measurement should represent.
Attribution is the expensive part, and it is where honest vendors get uncomfortable, correctly. Proving that a specific page change caused a specific citation gain is genuinely hard. There is no click path, no referrer in most cases, and no way to run a clean control.
The best available answer is not a claim of proof. It is a discipline: declared targets, repeated baselines, a prompt set fixed in advance, and a written account of what was changed and when, so correlation is at least legible rather than assembled after the fact.
Any vendor claiming causal proof in this category is overstating what the medium currently allows, and that overstatement is a better signal about the vendor than any case study.
What a buyer should require before signing
The prompt set, in writing, before work begins.
A baseline measured more than once, with the observed range stated rather than a single figure.
A target lift, named in advance, against that range.
A change log tying each production batch to a date, so any movement can be read against what actually shipped.
Those 4 requirements cost a vendor nothing if its method is real, and they are close to impossible to satisfy retroactively. That asymmetry is what makes them useful. An AEO agency engagement that agrees to all 4 up front has already distinguished itself from most of the market, before any work has happened.
Conclusion
Answer engine optimisation is a young discipline with a measurement culture borrowed from a more stable one, and the borrowing does not survive contact with how variable generated answers actually are.
The fix is not a better dashboard. It is committing to the prompt set, the baseline range and the target before the work, and accepting that what the category can honestly offer today is disciplined correlation rather than proof.
