Why does one answer fail as evidence?
A single good answer tells you it happened once, not that it happens. OpenAI's own documentation on ChatGPT search explains that ranking is based on a number of factors and there is "no way to guarantee top placement". It also says ChatGPT search "typically rewrites your query into one or more targeted queries" before it goes looking for results. That matters: the phrasing you typed is not the phrasing the system actually ran. Test one clever prompt and you have measured that prompt's luck, not your visibility. A prompt set spreads the risk across enough variation that a pattern can show through.
Google describes something structurally similar for its AI features: a "query fan-out" technique that lets these systems surface a wider and more diverse set of links than a single query would. Different platforms, same implication—the input you control is a starting point the system reworks, not the final query.
How do you build the prompt set?
Derive prompts from real buyer language, not from your own product vocabulary. Pull phrasing from sales calls, support tickets, review sites, and search terms you already know buyers use. Then cover the stages of a real buying journey rather than repeating one angle:
- Category discovery — "what are the best [service] providers for [need]."
- Comparison — "[option A] vs [option B] for [use case]."
- Best-for-Y framing — "best [service] for [specific outcome or constraint]."
- Local or qualified variants — adding a location, budget, or industry that narrows the field.
Keep the set small enough that you will actually repeat it—ten to fifteen prompts is a reasonable range. Then freeze it. Changing the wording between rounds breaks the comparison; you would no longer know whether an answer moved because the market moved or because you asked a different question.
Once a round has run, the prompt set is locked. New prompts start a new set, not a revision of the old one.
What controls keep the test honest?
Run every prompt from a temporary or logged-out session. OpenAI's memory documentation explains that ChatGPT can remember context from past chats to personalise responses, through saved memories and chat history, and that this can be turned off in Settings. Its Temporary Chat documentation confirms that Temporary Chats "do not use existing memories or create new memories" and will not appear in history. If you test while logged into an account that has spent months researching your own brand, you are measuring your own history, not the market's view. Temporary Chat, or a logged-out session, is the control condition.
Beyond that: note the model and the date for every run, since both are moving targets. And run each prompt more than once per round—outputs vary between runs on the same system, so a single pass through the list can mislead you either way.
It can tell you, for a fixed set of questions, whether your brand shows up, whether it is cited, whether the description is accurate, and whether that changes over time. It cannot tell you why an individual answer came out a particular way, cannot guarantee the next unscripted question behaves the same, and cannot prove causation between a change you shipped and a change in the answers—only correlation worth investigating further.
What should you log each round?
Track these as separate columns, because they are different achievements. Collapsing them into one score hides which problem you actually have.
Was the brand named at all, anywhere in the answer?
Was the site linked or cited as a source?
Was what it said about the business correct?
Did a competitor get named in the same slot instead?
A brand that is mentioned but never cited has a different problem to one that is cited but described incorrectly. Keep the columns apart so the next action is obvious.
Draft the prompt set from real buyer language and freeze it.
Run the full set under controlled conditions and log every column.
Repeat on a fixed schedule after changes ship—not daily.
Keep, revise, or stop based on the logged pattern, not one answer.
Re-run on a fixed cadence tied to real changes—after a site update, a new page, or a PR push—rather than checking every day. Daily checks mostly surface the variation between runs, and you end up reacting to noise.
What are the honest limits of this method?
This is observational, not experimental. Sample sizes are small by design, because the point is repeatability, not statistical power. You cannot attribute a change in the answers to any one action you took—too many other variables move at the same time, including the platform itself. Answers vary between runs of the same prompt; that variation is a known feature of how these systems sample responses, not a flaw in your method, and it is exactly why a fixed set beats one-off checks. And platforms change their retrieval and ranking behaviour without notice, so a prompt set that behaves one way this quarter is not guaranteed to behave the same way next quarter. Treat every result as a directional signal, not a verdict.
A prompt set does not prove why an answer changed. It proves whether it changed—across enough questions, and enough runs, to trust the pattern.
OpenAI Help Center: ChatGPT Search
OpenAI Help Center: Memory FAQ
OpenAI Help Center: Temporary Chat FAQ
Google Search Central: AI features and your website
Need the faster version first? Run the 60-second check to find a question worth tracking, then build the frozen prompt set around it.