Ask ChatGPT for “the best tools” today and your brand may appear. Ask again with a slightly different need, in another location, or through another interface and the answer may change. That does not necessarily mean the brand gained or lost visibility. It may mean you changed the experiment.
One query is a useful diagnostic. It is not a reliable measurement. A measurement must say what was tested, under which conditions, how many valid observations were collected, and how uncertainty was handled.
Why the same-looking query can produce a different answer
Generative outputs are variable by nature. OpenAI’s API documentation explicitly says model outputs can vary and recommends pinned model versions plus evaluations when consistent behavior matters. For ChatGPT search, the system may rewrite a user’s prompt into one or more targeted queries, and location can influence results. Conversation context can also affect whether search is used and how a follow-up is answered.
Google describes a related mechanism for AI features: “query fan-out” can issue multiple searches across subtopics and data sources. Google also notes that different models and techniques are used for different questions, so the answers and supporting links can vary.
If any factor changes, you may be looking at a different test. Even when all recorded inputs appear identical, sampling and hidden platform changes mean an answer should be treated as an observation—not a permanent property of the brand.
How a one-query score creates false confidence
Suppose a brand appears in one answer. Calling that a “100% mention rate” is mathematically true only for that one valid observation, but it invites a much larger interpretation than the data supports. The reverse is equally dangerous: absence from one answer does not prove zero visibility.
| Result | Defensible conclusion | Conclusion you cannot make |
|---|---|---|
| Brand appears once | The brand was mentioned for this prompt in this observed answer. | “The brand has 100% ChatGPT visibility.” |
| Brand is absent once | The brand was not mentioned in this observed answer. | “ChatGPT never recommends the brand.” |
| Request fails | No valid observation was collected. | “The brand scored zero.” |
| A competitor appears | That competitor was present in this answer. | “It is the category leader” without repeated, comparable tests. |
Aivius field evidence: one prompt did not have one universal winner
The Aivius 2026 AI Video Enhancement Benchmark asked the same 50 English prompts to ChatGPT, Gemini, and Perplexity on August 23, 2026. Across the 14 highest-intent prompts, seven did not produce the same primary pick on all three engines.
| Matched prompt | ChatGPT primary pick | Gemini primary pick | Perplexity primary pick |
|---|---|---|---|
| Topaz vs AVCLabs | Topaz Video AI | Topaz Video AI | AVCLabs |
| HitPaw vs Topaz | HitPaw VikPea | Topaz Video AI | Topaz Video AI |
| best Topaz alternative | Aiarty | Aiarty | Aiarty |
The first two rows demonstrate platform disagreement under a matched prompt; the third shows that agreement can also occur. A single ChatGPT answer therefore cannot be generalized to “AI visibility” across engines.
A repeatable AI visibility measurement protocol
Define the decision before collecting data
Are you diagnosing a page, comparing competitors, tracking a campaign, or estimating category share of voice? The decision determines the prompts, sample size, cadence, and acceptable uncertainty.
Build a prompt set from real buyer intents
Group prompts into discovery, problem/solution, comparison, validation, and purchase intent. Keep a canonical version of each prompt. A long list of near-duplicates can look like a large sample while measuring the same question repeatedly.
Freeze the observable test conditions
Record the exact prompt, platform, interface or API, model identifier when available, search/web state, language, location, clean or continued conversation, and timestamp. Do not silently mix ChatGPT’s consumer interface with an API model.
Repeat each measurement point
Run more than once when the decision depends on stability. Repetitions help distinguish a persistent pattern from answer-level variation. There is no universal magic number: use more repetitions for close competitor comparisons, high-stakes decisions, or unstable prompts.
Separate valid absence from technical failure
A valid answer with no brand mention belongs in the denominator. A timeout, quota error, missing API key, or parse failure does not. Report platform completion and failure reasons beside the score.
Classify evidence consistently
Use explicit rules for mention, recommendation, and citation. Preserve the answer and source URLs so a reviewer can audit borderline cases.
Aggregate at the right level
Calculate rates by prompt group and platform before producing an overall number. Weighting should reflect business importance, not whichever group happened to contain more prompts. Always display numerator, denominator, and completion rate.
Compare matched time windows
Use the same prompt set and conditions for trend comparisons. Document model or protocol changes. A before/after result is not interpretable if the measurement instrument changed between periods.
How many prompts and repetitions do you need?
There is no honest single answer without a target level of precision and a model of the data. A practical pilot can begin with a small, deliberately balanced set, but it must be labeled accordingly. Expand the sample after inspecting where results vary.
These are descriptive labels, not statistical certifications. Aivius recommends attaching a confidence label based on sample coverage, repetitions, platform completion, and classification quality. Do not infer “high confidence” merely from a large prompt count if all prompts are near-duplicates.
A practical starting design
- choose 4–6 commercially meaningful topic groups;
- include several distinct prompts per group;
- repeat each prompt under matched conditions;
- run platforms independently and show completion per platform;
- retain raw evidence long enough to audit changes;
- review the prompt set periodically, but version changes instead of rewriting history.
A free scan with a small sample can still be valuable. Its job is to find signals worth investigating, not to claim a precise market share. Label it “indicative,” show the number of valid samples, and let users inspect why a platform failed or why a brand was classified as absent.
What a trustworthy dashboard should show
| Field | Why it matters |
|---|---|
| Valid observations / scheduled observations | Exposes missing data instead of converting failures to zero. |
| Prompt and topic group | Makes the business intent and weighting auditable. |
| Platform and interface | Prevents unlike systems from being blended invisibly. |
| Matched text and source URL | Lets a human verify mentions and citations. |
| Timestamp, language, and location | Preserves conditions that can affect retrieval and output. |
| Confidence label and protocol version | Communicates limits and protects trend comparability. |
Sources and further reading
- OpenAI API: backward compatibility—output variability, pinned model versions, and evaluations.
- OpenAI Help: Searching the web with ChatGPT—query rewriting, location, citations, and verification cautions.
- OpenAI: Introducing ChatGPT search—search behavior, follow-up context, and source links.
- Google Search Central: AI features and your website—query fan-out and varying models, answers, and links.
- Aivius: High-intent query optimization.
Start with an honest snapshot
Aivius reports valid samples and platform completion separately, so an API failure is never presented as a visibility loss.