Generative engine optimization is often reduced to a single question: “Did the AI mention us?” That question is useful, but incomplete. A name appearing in an answer does not prove that the system prefers the brand, trusts the brand’s website, or will help a buyer choose it.

A more useful measurement model separates three observable events: mention, recommendation, and citation. They belong to different parts of the buyer journey and should be recorded independently.

Scope: The definitions below are Aivius operational rules for repeatable measurement. They are not official definitions published by ChatGPT, Perplexity, or Google.
Signal 1

Mention

The target brand, product, domain, or verified alias appears in the answer text. No link or positive evaluation is required.

Signal 2

Recommendation

The answer explicitly presents the brand as a candidate, preferred option, or suitable choice for the user’s decision.

Signal 3

Citation

The answer contains an attributable, parsable source reference such as a URL. The cited source and mentioned brand are stored separately.

Why the distinction changes business decisions

These signals answer different questions. Mentions show whether a brand is present in the answer set. Recommendations show whether it enters the decision set. Citations show which sources the system exposes as evidence. Combining them into one binary “visibility” flag destroys that information.

Observed answerMentionRecommendationCitation
“Acme launched a desktop editor in 2025.”YesNoNo
“For small teams that need offline editing, Acme is a strong option.”YesYesNo
“The format supports 10-bit color,” linked to Acme’s documentation.Possibly notNoYes: Acme source

The examples are illustrative, not evidence about a real brand. The third row highlights a common analytics error: an AI can use a company’s page as evidence without naming the company in its prose. Conversely, it can recommend a product while citing a publisher, marketplace, or review site instead of the product’s domain.

Aivius field evidence: citations and recommendations diverged

Aivius’s 2026 AI Video Enhancement Benchmark provides a real counterexample to treating citations as recommendations. On August 23, 2026, Aivius collected 50 English prompts across ChatGPT, Gemini, and Perplexity—150 valid answers in total. Recommendation counts came from all three engines; the owned-domain citation counts below came from source URLs returned with the 50 Perplexity answers.

BrandRecommendations across 150 answersOwned-domain citations in 50 Perplexity answersObserved pattern
Topaz Video AI6222High recommendation coverage and strong first-party retrieval
UniFab922Same owned-domain citation count as Topaz, far fewer recommendations
VideoProc Converter AI1013First-party retrieval present, limited recommendation conversion
What the data supports: citation presence and recommendation frequency were not interchangeable in this sample. What it does not support: citations causing—or failing to cause—recommendations. The benchmark is a dated, single-run, English-language category sample.

The practical implication is not “citations do not matter.” It is that teams need two separate questions: Is our evidence being retrieved? and Is our brand being selected for the user’s decision? UniFab’s pattern would be invisible in a dashboard that collapsed both into one score.

1. What counts as an AI mention?

A mention is the narrowest signal. Count it when the answer includes an exact brand or product name, the canonical domain, or a pre-approved alias. Do not use unrestricted fuzzy matching: a generic word that happens to resemble a brand can create false positives.

Store more than a yes/no field

  • the exact matched text and the surrounding sentence;
  • the entity it maps to—brand, product, parent company, or domain;
  • position in a list, if the answer provides an ordered list;
  • sentiment or context, kept separate from mention status;
  • the prompt, platform, model or interface, language, location, and timestamp.

A negative mention is still a mention. It should not be converted into a positive visibility outcome. This is why a useful dashboard shows presence and context as different dimensions.

2. What counts as a recommendation?

A recommendation requires decision language, not mere description. The answer must position the brand as suitable for a need, include it among options the user should consider, or prefer it under stated conditions.

“Acme has an API” is a factual statement. “Choose Acme if you need an API with local processing” is a recommendation. A ranked list is not automatically an endorsement unless the wording or requested task makes the list a set of recommended choices.

Do not infer intent from the brand name alone. Recommendation classification should consider the prompt and the full answer context. Where the wording is ambiguous, store “uncertain” and route a sample to human review.

3. What counts as a citation?

A citation is an explicit source relationship exposed in the response. ChatGPT search answers can include inline citations and a Sources panel; OpenAI also warns that citations may be incomplete or incorrect, so important claims should be checked against the underlying source. Perplexity’s current API exposes source metadata through its search_results field. Google says AI Overviews and AI Mode can surface supporting web links, while eligibility still depends on the ordinary requirements for appearing in Google Search.

For measurement, parse the destination URL, normalize the domain, preserve the original URL, and connect the citation to the claim or answer segment where possible. A page being crawlable or eligible is not evidence that it was cited in a specific answer.

The fourth state dashboards often hide: unknown

A failed request is not a zero. If a platform times out, a quota is exhausted, a key is missing, or a response cannot be parsed, the correct state is unknown or unmeasured. Treating it as “not mentioned” lowers the score for an operational failure rather than an observed market outcome.

StatusWhat it meansHow it affects rates
Observed: yesThe signal was found in a valid answer.Include in numerator and denominator.
Observed: noA valid answer was checked and the signal was absent.Include in denominator.
UnknownThe platform or parser did not produce a valid observation.Exclude from the rate; report separately.

A minimum viable AI visibility scorecard

  1. Mention rate: valid answers that mention the entity ÷ valid answers checked.
  2. Recommendation rate: valid answers that recommend the entity ÷ valid decision-intent answers checked.
  3. Owned citation rate: valid answers citing an owned domain ÷ valid answers checked.
  4. Third-party citation rate: valid answers citing independent coverage about the entity ÷ valid answers checked.
  5. Completion rate: valid observations ÷ scheduled observations. Never conceal this behind the visibility score.

Report raw counts beside percentages—“4 of 12 valid answers,” not only “33%.” Then segment by platform, topic, and buyer intent. A blended score can be useful for trend summaries, but it should never replace the underlying evidence.

How to implement the framework

  1. Create an entity dictionary. Record the canonical brand, products, domains, parent entity, and safe aliases.
  2. Define query groups. Separate discovery, comparison, problem-solving, and purchase-intent prompts.
  3. Capture the evidence. Store the prompt, answer, sources, request status, and reproducibility metadata.
  4. Classify each signal independently. A single answer can contain any combination of mention, recommendation, and citation.
  5. Audit samples. Review ambiguous classifications and track disagreement between automated and human labels.
  6. Measure trends, not anecdotes. Use the repeatable protocol in our guide to why one ChatGPT query is not enough.

Sources and further reading

Measure evidence, not vanity signals

Run an indicative scan across AI platforms, then inspect mentions, citations, failures, and sample size separately.