Measurement guide
How to measure AI citations and recommendations
By Brandilite · Guide version 11 September 2026
AI visibility measurement starts with a fixed set of buyer questions and saved answers. Record whether an answer mentions your brand, recommends it and cites your website as separate outcomes. Repeat the same questions under documented conditions, then report the counts behind each percentage.
This guide sets out Brandilite's proposed observation method. It is a practical protocol, not an industry standard or a report of completed results. The worked example below is illustrative. Use it to build a baseline before setting improvement targets.
Define what counts before collecting answers
An observation is one submitted prompt and its resulting answer on one platform at one time. Score each outcome once per answer:
- Brand mention: the answer text explicitly names the brand. A domain appearing only in a source card does not meet this text-mention definition.
- Recommendation: the answer affirmatively presents the brand as a suitable provider for the requested need. A negative comparison, passing reference or instruction to research it further is not automatically a recommendation.
- Owned citation: a visible citation or supporting source link points to an agreed brand-owned domain. A bare URL in a list of providers is a navigational link unless the answer presents it as a source.
- Third-party support: an external source is cited for a statement about the brand, and its page supports that statement. Save the relevant passage; an unrelated link from a respected publisher is insufficient.
These outcomes can overlap. A guide can receive a citation without a provider recommendation. Record factual errors separately: a visible citation can still support an inaccurate answer. Do not collapse the four measures into an unexplained score.
Freeze a relevant prompt panel
Start with 20 questions grounded in your service, customers and buying process. Separate recommendation tests, such as “Which AEO agencies serve US B2B SaaS companies?”, from citation tests, such as “How should a SaaS company measure citations in AI answers?” This protocol uses 14 recommendation prompts and six citation prompts.
Give each prompt an ID, exact wording, audience, intent and version. Keep branded checks, such as “What does Brandilite do?”, outside unbranded discovery rates. Naming your company in the question changes the test.
Use customer questions and search research to choose the panel. Search volume can help prioritize topics, but it does not measure how frequently people submit an AI prompt. Agree on the panel before seeing results; document additions as a new version. Our prompt research page explains the distinction between question design and measured answers.
Run a repeatable collection wave
Use ChatGPT Search, Perplexity Search and Google AI Mode where available. Record the actual interface and model or mode; use “not exposed” when the model is undisclosed. Keep consumer interfaces and API runs in separate datasets.
Run every prompt three times on separate days within a defined baseline window, for example days 1, 3 and 5. Twenty prompts across three platforms and three repetitions produce 180 planned observations. Repeated answers show variability within this sample; they are not 180 independent buyers.
Start each observation in a fresh conversation. Use the same account context where practical and record whether memory or personalization is enabled. Fresh conversations do not remove all personalization. Record language, device, actual location settings, timestamp in UTC and whether web search was used. US wording alone does not establish US geolocation: ChatGPT can use approximate location and relevant saved memories. OpenAI's search documentation
Do not coach a weak answer toward your brand. Save it before any follow-up. Repeat the complete wave around days 30, 60 and 90 using equivalent collection windows. Ten fixed priority prompts, run once per platform weekly, provide 30 directional checks between full waves.
Preserve evidence another analyst can review
Save the exact prompt, complete answer, timestamp, source links and a screenshot or export under the run ID. Keep account identifiers and private conversations in controlled storage; a public share link is unnecessary.
Open the citations and preserve their destination URLs. Retain the original URL alongside any normalized reporting URL; strip tracking parameters only when they do not change the content. Deduplicate repeated links for a separate unique-page count, while counting an owned citation only once per answer.
Distinguish sources supporting the answer from additional suggested links. ChatGPT's Sources view can include both cited sources and other relevant links; Perplexity provides linked sources for verification. Inspect the passage and destination rather than relying on a domain logo. OpenAI source guidance, Perplexity's explanation
Have a second reviewer check borderline recommendations and unsupported claims. Record disagreements and their resolution. A cited URL that no longer opens remains an observed citation, with its verification status recorded as inaccessible.
Keep failures out of visibility denominators
Use Planned, Completed, Failed or Unavailable as the run status. Completed means a usable answer was captured and reviewed. For completed runs, record Yes or No for each scored outcome. Leave outcomes blank for unrun or failed observations and write the reason.
An accessible answer that omits your brand is a valid No. A timeout, login barrier or unavailable search mode is not. Keep failed attempts in the log; a retry receives a new attempt ID linked to the planned run, with only one accepted completed answer per planned observation.
If studying Google AI Overviews separately, record whether an overview appeared. Google says they do not always trigger. Report the trigger rate and citation rate among triggered, reviewable overviews separately. Do not substitute an ordinary search listing for an AI answer. Google's AI features guidance
Calculate rates with visible denominators
Calculate each metric by platform and prompt cohort first. Include only completed answers with evidence and a reviewed Yes/No value for that metric. Flag missing scores for review instead of silently treating blanks as No.
- Mention rate = completed eligible answers with a brand mention ÷ completed eligible answers scored for mentions × 100.
- Recommendation rate = completed recommendation-test answers recommending the brand ÷ completed recommendation-test answers scored for recommendations × 100.
- Owned citation rate = completed eligible answers citing an owned domain ÷ completed eligible answers scored for owned citations × 100. Report recommendation and citation cohorts separately.
- Completion rate = accepted completed observations ÷ planned observations × 100.
Show the numerator, denominator, planned count and collection dates beside every rate. When the denominator is zero, report “unavailable.” For pooled reporting, disclose the platform mix and use total successes divided by total eligible answers. Compare matching panels and settings; changed coverage can move a pooled rate without any improvement.
Worked example: illustrative observations only
Suppose one platform has 12 planned recommendation-test observations. Ten complete with saved evidence and two fail. Among the ten completed answers, six mention ExampleCo, four recommend it, three cite its website and two contain verified third-party support.
| Measure | Calculation | Illustrative result |
|---|---|---|
| Completion | 10 / 12 | 83.3% |
| Brand mention | 6 / 10 | 60% |
| Recommendation | 4 / 10 | 40% |
| Owned citation | 3 / 10 | 30% |
| Verified third-party support | 2 / 10 | 20% |
The categories overlap; adding their percentages is meaningless. The two failures remain visible but do not become negative answers. These invented counts demonstrate the arithmetic only. They are not Brandilite results, a market benchmark or evidence of improvement.
Connect visibility to business outcomes carefully
Keep answer observations, search performance, visits and conversions as distinct datasets. Google includes AI-feature traffic within Search Console's Web performance reporting; those totals are not your sampled citation denominator. Brandilite's search-evidence index methodology describes a dated proxy, not direct assistant recommendations. Google measurement guidance
In GA4, review observed session source/medium and landing pages. Maintain an explicit list of AI referral sources based on actual data. Direct visits can have missing referral information; do not relabel all direct traffic as AI. Separate session acquisition from attribution credit, which depends on the reporting scope and attribution model. Google's traffic-source definitions
Define conversions at verified outcomes: a successfully accepted inquiry can trigger generate_lead; a completed payment can trigger purchase. A CTA click, form opening or audit request is not a purchase. Deduplicate events and validate them before marking appropriate events as key events. Keep lead qualification in the CRM or agreed business workflow. GA4 recommended events
Respect the site's consent implementation and visitor choices. Consent mode controls tag behavior; it does not collect consent for you. Test denied and granted states, and document measurement gaps. Never send names, emails, raw inquiry text or private answer transcripts to GA4. Google consent guidance, Google's PII restrictions
An observed AI referral followed by a qualified inquiry is useful evidence. It does not prove that a particular sampled citation caused the inquiry, or that all AI-assisted journeys were captured. Report optional self-reported discovery answers separately from measured referrals.
Start your observation sheet
Use one row per attempt. Copy the header below into a CSV file or spreadsheet, leaving all outcomes blank until an answer is reviewed. Keep the prompt panel and this scoring rubric alongside it. Store detailed answer evidence separately and link it through the run ID.
"planned_run_id","attempt_id","accepted_attempt","wave","panel_version","prompt_id","test_type","exact_prompt","engine","actual_model_mode","repeat","run_at_utc","language","device","actual_country_setting","account_context","personalization_context","fresh_conversation","search_used","status","overview_triggered","brand_mention","brand_recommended","owned_citation","third_party_support","owned_citation_urls_original","owned_citation_urls_normalized","third_party_support_url","third_party_support_passage","answer_evidence_file_or_url","citation_verification_status","factual_errors","failure_reason","reviewer","reviewed_at_utc","review_notes"
"","","","","","","","","","","","","","","","","","","","","","","","","","","","","","","","","","","",""
Before each review, reconcile planned, completed, failed and unavailable counts; inspect citation destinations; resolve ambiguous scores; and list the page improvements supported by the evidence. Explore Brandilite's AI visibility service or review the audit workflow to see how measurement informs implementation priorities.