Ask ChatGPT for “the best tools” today and your brand may appear. Ask again with a slightly different need, in another location, or through another interface and the answer may change. That does not necessarily mean the brand gained or lost visibility. It may mean you changed the experiment.

One query is a useful diagnostic. It is not a reliable measurement. A measurement must say what was tested, under which conditions, how many valid observations were collected, and how uncertainty was handled.

Core rule: never turn one generated answer into a market-wide percentage. Report it as one observation, preserve the evidence, and label the result as indicative until the sample is large and repeatable enough for its intended decision.

Why the same-looking query can produce a different answer

Generative outputs are variable by nature. OpenAI’s API documentation explicitly says model outputs can vary and recommends pinned model versions plus evaluations when consistent behavior matters. For ChatGPT search, the system may rewrite a user’s prompt into one or more targeted queries, and location can influence results. Conversation context can also affect whether search is used and how a follow-up is answered.

Google describes a related mechanism for AI features: “query fan-out” can issue multiple searches across subtopics and data sources. Google also notes that different models and techniques are used for different questions, so the answers and supporting links can vary.

Promptwording, intent, constraints
Systemplatform, model, search state
Contexthistory, language, location
Timefresh sources and model changes

If any factor changes, you may be looking at a different test. Even when all recorded inputs appear identical, sampling and hidden platform changes mean an answer should be treated as an observation—not a permanent property of the brand.

How a one-query score creates false confidence

Suppose a brand appears in one answer. Calling that a “100% mention rate” is mathematically true only for that one valid observation, but it invites a much larger interpretation than the data supports. The reverse is equally dangerous: absence from one answer does not prove zero visibility.

ResultDefensible conclusionConclusion you cannot make
Brand appears onceThe brand was mentioned for this prompt in this observed answer.“The brand has 100% ChatGPT visibility.”
Brand is absent onceThe brand was not mentioned in this observed answer.“ChatGPT never recommends the brand.”
Request failsNo valid observation was collected.“The brand scored zero.”
A competitor appearsThat competitor was present in this answer.“It is the category leader” without repeated, comparable tests.

Aivius field evidence: one prompt did not have one universal winner

The Aivius 2026 AI Video Enhancement Benchmark asked the same 50 English prompts to ChatGPT, Gemini, and Perplexity on August 23, 2026. Across the 14 highest-intent prompts, seven did not produce the same primary pick on all three engines.

Matched promptChatGPT primary pickGemini primary pickPerplexity primary pick
Topaz vs AVCLabsTopaz Video AITopaz Video AIAVCLabs
HitPaw vs TopazHitPaw VikPeaTopaz Video AITopaz Video AI
best Topaz alternativeAiartyAiartyAiarty

The first two rows demonstrate platform disagreement under a matched prompt; the third shows that agreement can also occur. A single ChatGPT answer therefore cannot be generalized to “AI visibility” across engines.

Important limit: this benchmark was one dated run per engine. It demonstrates cross-platform variation, not run-to-run instability within the same platform. Aivius is separately collecting matched repeat runs; those results will not be published until at least three complete, auditable observations are available.

A repeatable AI visibility measurement protocol

  1. Define the decision before collecting data

    Are you diagnosing a page, comparing competitors, tracking a campaign, or estimating category share of voice? The decision determines the prompts, sample size, cadence, and acceptable uncertainty.

  2. Build a prompt set from real buyer intents

    Group prompts into discovery, problem/solution, comparison, validation, and purchase intent. Keep a canonical version of each prompt. A long list of near-duplicates can look like a large sample while measuring the same question repeatedly.

  3. Freeze the observable test conditions

    Record the exact prompt, platform, interface or API, model identifier when available, search/web state, language, location, clean or continued conversation, and timestamp. Do not silently mix ChatGPT’s consumer interface with an API model.

  4. Repeat each measurement point

    Run more than once when the decision depends on stability. Repetitions help distinguish a persistent pattern from answer-level variation. There is no universal magic number: use more repetitions for close competitor comparisons, high-stakes decisions, or unstable prompts.

  5. Separate valid absence from technical failure

    A valid answer with no brand mention belongs in the denominator. A timeout, quota error, missing API key, or parse failure does not. Report platform completion and failure reasons beside the score.

  6. Classify evidence consistently

    Use explicit rules for mention, recommendation, and citation. Preserve the answer and source URLs so a reviewer can audit borderline cases.

  7. Aggregate at the right level

    Calculate rates by prompt group and platform before producing an overall number. Weighting should reflect business importance, not whichever group happened to contain more prompts. Always display numerator, denominator, and completion rate.

  8. Compare matched time windows

    Use the same prompt set and conditions for trend comparisons. Document model or protocol changes. A before/after result is not interpretable if the measurement instrument changed between periods.

How many prompts and repetitions do you need?

There is no honest single answer without a target level of precision and a model of the data. A practical pilot can begin with a small, deliberately balanced set, but it must be labeled accordingly. Expand the sample after inspecting where results vary.

Indicative: small or incomplete sampleDirectional: repeated, balanced sampleDecision-grade: stable protocol and audited evidence

These are descriptive labels, not statistical certifications. Aivius recommends attaching a confidence label based on sample coverage, repetitions, platform completion, and classification quality. Do not infer “high confidence” merely from a large prompt count if all prompts are near-duplicates.

A practical starting design

  • choose 4–6 commercially meaningful topic groups;
  • include several distinct prompts per group;
  • repeat each prompt under matched conditions;
  • run platforms independently and show completion per platform;
  • retain raw evidence long enough to audit changes;
  • review the prompt set periodically, but version changes instead of rewriting history.

A free scan with a small sample can still be valuable. Its job is to find signals worth investigating, not to claim a precise market share. Label it “indicative,” show the number of valid samples, and let users inspect why a platform failed or why a brand was classified as absent.

What a trustworthy dashboard should show

FieldWhy it matters
Valid observations / scheduled observationsExposes missing data instead of converting failures to zero.
Prompt and topic groupMakes the business intent and weighting auditable.
Platform and interfacePrevents unlike systems from being blended invisibly.
Matched text and source URLLets a human verify mentions and citations.
Timestamp, language, and locationPreserves conditions that can affect retrieval and output.
Confidence label and protocol versionCommunicates limits and protects trend comparability.

Sources and further reading

Start with an honest snapshot

Aivius reports valid samples and platform completion separately, so an API failure is never presented as a visibility loss.