Measurement guide

How to measure AI citations and recommendations

By Brandilite · Guide version 11 September 2026

AI visibility measurement starts with a fixed set of buyer questions and saved answers. Record whether an answer mentions your brand, recommends it and cites your website as separate outcomes. Repeat the same questions under documented conditions, then report the counts behind each percentage.

This guide sets out Brandilite's proposed observation method. It is a practical protocol, not an industry standard or a report of completed results. The worked example below is illustrative. Use it to build a baseline before setting improvement targets.

Define what counts before collecting answers

An observation is one submitted prompt and its resulting answer on one platform at one time. Score each outcome once per answer:

These outcomes can overlap. A guide can receive a citation without a provider recommendation. Record factual errors separately: a visible citation can still support an inaccurate answer. Do not collapse the four measures into an unexplained score.

Freeze a relevant prompt panel

Start with 20 questions grounded in your service, customers and buying process. Separate recommendation tests, such as “Which AEO agencies serve US B2B SaaS companies?”, from citation tests, such as “How should a SaaS company measure citations in AI answers?” This protocol uses 14 recommendation prompts and six citation prompts.

Give each prompt an ID, exact wording, audience, intent and version. Keep branded checks, such as “What does Brandilite do?”, outside unbranded discovery rates. Naming your company in the question changes the test.

Use customer questions and search research to choose the panel. Search volume can help prioritize topics, but it does not measure how frequently people submit an AI prompt. Agree on the panel before seeing results; document additions as a new version. Our prompt research page explains the distinction between question design and measured answers.

Run a repeatable collection wave

Use ChatGPT Search, Perplexity Search and Google AI Mode where available. Record the actual interface and model or mode; use “not exposed” when the model is undisclosed. Keep consumer interfaces and API runs in separate datasets.

Run every prompt three times on separate days within a defined baseline window, for example days 1, 3 and 5. Twenty prompts across three platforms and three repetitions produce 180 planned observations. Repeated answers show variability within this sample; they are not 180 independent buyers.

Start each observation in a fresh conversation. Use the same account context where practical and record whether memory or personalization is enabled. Fresh conversations do not remove all personalization. Record language, device, actual location settings, timestamp in UTC and whether web search was used. US wording alone does not establish US geolocation: ChatGPT can use approximate location and relevant saved memories. OpenAI's search documentation

Do not coach a weak answer toward your brand. Save it before any follow-up. Repeat the complete wave around days 30, 60 and 90 using equivalent collection windows. Ten fixed priority prompts, run once per platform weekly, provide 30 directional checks between full waves.

Preserve evidence another analyst can review

Save the exact prompt, complete answer, timestamp, source links and a screenshot or export under the run ID. Keep account identifiers and private conversations in controlled storage; a public share link is unnecessary.

Open the citations and preserve their destination URLs. Retain the original URL alongside any normalized reporting URL; strip tracking parameters only when they do not change the content. Deduplicate repeated links for a separate unique-page count, while counting an owned citation only once per answer.

Distinguish sources supporting the answer from additional suggested links. ChatGPT's Sources view can include both cited sources and other relevant links; Perplexity provides linked sources for verification. Inspect the passage and destination rather than relying on a domain logo. OpenAI source guidance, Perplexity's explanation

Have a second reviewer check borderline recommendations and unsupported claims. Record disagreements and their resolution. A cited URL that no longer opens remains an observed citation, with its verification status recorded as inaccessible.

Keep failures out of visibility denominators

Use Planned, Completed, Failed or Unavailable as the run status. Completed means a usable answer was captured and reviewed. For completed runs, record Yes or No for each scored outcome. Leave outcomes blank for unrun or failed observations and write the reason.

An accessible answer that omits your brand is a valid No. A timeout, login barrier or unavailable search mode is not. Keep failed attempts in the log; a retry receives a new attempt ID linked to the planned run, with only one accepted completed answer per planned observation.

If studying Google AI Overviews separately, record whether an overview appeared. Google says they do not always trigger. Report the trigger rate and citation rate among triggered, reviewable overviews separately. Do not substitute an ordinary search listing for an AI answer. Google's AI features guidance

Calculate rates with visible denominators

Calculate each metric by platform and prompt cohort first. Include only completed answers with evidence and a reviewed Yes/No value for that metric. Flag missing scores for review instead of silently treating blanks as No.

Show the numerator, denominator, planned count and collection dates beside every rate. When the denominator is zero, report “unavailable.” For pooled reporting, disclose the platform mix and use total successes divided by total eligible answers. Compare matching panels and settings; changed coverage can move a pooled rate without any improvement.

Worked example: illustrative observations only

Suppose one platform has 12 planned recommendation-test observations. Ten complete with saved evidence and two fail. Among the ten completed answers, six mention ExampleCo, four recommend it, three cite its website and two contain verified third-party support.

Illustrative arithmetic only — not Brandilite results
MeasureCalculationIllustrative result
Completion10 / 1283.3%
Brand mention6 / 1060%
Recommendation4 / 1040%
Owned citation3 / 1030%
Verified third-party support2 / 1020%

The categories overlap; adding their percentages is meaningless. The two failures remain visible but do not become negative answers. These invented counts demonstrate the arithmetic only. They are not Brandilite results, a market benchmark or evidence of improvement.

Connect visibility to business outcomes carefully

Keep answer observations, search performance, visits and conversions as distinct datasets. Google includes AI-feature traffic within Search Console's Web performance reporting; those totals are not your sampled citation denominator. Brandilite's search-evidence index methodology describes a dated proxy, not direct assistant recommendations. Google measurement guidance

In GA4, review observed session source/medium and landing pages. Maintain an explicit list of AI referral sources based on actual data. Direct visits can have missing referral information; do not relabel all direct traffic as AI. Separate session acquisition from attribution credit, which depends on the reporting scope and attribution model. Google's traffic-source definitions

Define conversions at verified outcomes: a successfully accepted inquiry can trigger generate_lead; a completed payment can trigger purchase. A CTA click, form opening or audit request is not a purchase. Deduplicate events and validate them before marking appropriate events as key events. Keep lead qualification in the CRM or agreed business workflow. GA4 recommended events

Respect the site's consent implementation and visitor choices. Consent mode controls tag behavior; it does not collect consent for you. Test denied and granted states, and document measurement gaps. Never send names, emails, raw inquiry text or private answer transcripts to GA4. Google consent guidance, Google's PII restrictions

An observed AI referral followed by a qualified inquiry is useful evidence. It does not prove that a particular sampled citation caused the inquiry, or that all AI-assisted journeys were captured. Report optional self-reported discovery answers separately from measured referrals.

Start your observation sheet

Use one row per attempt. Copy the header below into a CSV file or spreadsheet, leaving all outcomes blank until an answer is reviewed. Keep the prompt panel and this scoring rubric alongside it. Store detailed answer evidence separately and link it through the run ID.

"planned_run_id","attempt_id","accepted_attempt","wave","panel_version","prompt_id","test_type","exact_prompt","engine","actual_model_mode","repeat","run_at_utc","language","device","actual_country_setting","account_context","personalization_context","fresh_conversation","search_used","status","overview_triggered","brand_mention","brand_recommended","owned_citation","third_party_support","owned_citation_urls_original","owned_citation_urls_normalized","third_party_support_url","third_party_support_passage","answer_evidence_file_or_url","citation_verification_status","factual_errors","failure_reason","reviewer","reviewed_at_utc","review_notes"
"","","","","","","","","","","","","","","","","","","","","","","","","","","","","","","","","","","",""

Before each review, reconcile planned, completed, failed and unavailable counts; inspect citation destinations; resolve ambiguous scores; and list the page improvements supported by the evidence. Explore Brandilite's AI visibility service or review the audit workflow to see how measurement informs implementation priorities.