Build a fixed prompt set first
Write 20-40 prompts a real buyer would type, spanning discovery ('best tools for X'), comparison ('X vs Y'), and validation ('is X any good'). Freeze the wording. Everything else in this workflow depends on running the same prompts repeatedly.
Measure with browsing enabled
Run the prompt set with web access on and log, per prompt: whether you were mentioned, in what position, which of your URLs was cited, and which competitors appeared. Repeat three times per prompt - responses vary, so treat a single run as noise.
Diagnose the misses
For each prompt where you were absent, ask the model what sources it used and what would have to be true of a source to be included. The answer is not authoritative, but it reliably surfaces the format and evidence gap: missing pricing table, no comparison page, no dated benchmark.
Draft answer blocks, then rewrite them
Ask for a 50-word direct answer to each target question, then edit hard: add your own numbers, remove hedging, and delete anything the model could have produced without your data. Unedited AI copy is exactly the generic text that synthesis discards.
Close entity gaps
Ask ChatGPT to describe your company from memory. Whatever it gets wrong or omits is a public-record gap. Fix it where the facts live - your About page, organization schema, directory profiles, and third-party coverage - not by repeating it in a blog post.
Re-run and log the delta
Four to six weeks after shipping fixes, run the frozen prompt set again and compare mention rate and citation rate. That delta is your AEO/GEO progress report, and it is the only one that reflects the surface your buyers use.