AI search visibility is best measured as a set of separate observations: mentions, citations, citation support, factual accuracy, engine variance, referral sessions, and business outcomes. Keep the prompt set stable, document every run, inspect cited pages manually, and report sample limits instead of turning a changing answer set into a false universal score.
- Track mentions, citations, factual support, referral sessions, and business outcomes separately. They answer different questions and should not be combined into one score.
- Use a documented prompt set and record the engine, account state, location, date, and model where available. Generated answers vary between runs.
- Review cited pages manually. A citation is useful only when the link resolves and the page supports the claim made in the answer.
- Treat changes as observations from your sample. Do not promise ranking, citation, traffic, or revenue gains from a visibility score.
Build a Prompt Set You Can Reuse

Start with the questions a buyer might ask before knowing your brand. Group them by intent, such as category research, problem solving, comparisons, implementation, and direct brand checks. Keep branded prompts in their own group so they do not inflate the unbranded result.
There is no universal mix of prompt types. A local service company, a software vendor, and an online shop have different buying journeys. Document why each prompt is included, who it represents, and which market or language it covers.
Freeze the set for a reporting period. If you add or remove prompts, show the change instead of comparing the new score directly with the old one.
For every run, retain the exact prompt, returned answer, cited URLs, engine, date, location, account state, and model or product mode when the interface exposes it. Run the same set more than once when the budget allows. Assistant answers can change between runs, so one answer is an observation, not a stable ranking.
Use a simple diagnostic sequence:
- Was the brand mentioned?
- Was its own domain cited?
- Did the cited page support the answer?
- Were the stated brand facts correct and current?
- Did a measurable visit or conversion follow?
Each question points to a different problem. Low mention coverage may indicate weak category relevance. Mentions without citations call for a source review. Broken or unsupported citations require page and claim checks. Traffic without useful outcomes is an analytics or conversion issue.
Use Metrics With Clear Operational Definitions
A practical report does not need fourteen overlapping scores. Start with metrics that your team can calculate and audit:
- Prompt coverage: valid runs completed divided by planned runs.
- Brand mention rate: valid answers that mention the brand divided by valid answers in the same prompt group.
- Domain citation rate: valid answers that cite the brand's domain divided by valid answers in the same group.
- Supported citation rate: reviewed citations whose destination page supports the associated claim divided by citations reviewed.
- Factual accuracy rate: checked brand claims scored correct and current divided by all checked claims.
- Cited URL mix: the pages and page types that appear in citations.
- Engine variance: the difference between comparable observations from each engine. Keep this descriptive unless the sample and method justify statistical analysis.
- AI referral sessions: visits attributed to an assistant or AI search source in analytics. Some visits will have missing or incomplete referral data.
- Assisted outcomes: leads, signups, purchases, or other selected events from attributed sessions. State the attribution window and model.
Keep the denominators visible. A citation rate based on six usable answers should not look as reliable as one based on a larger, repeated sample. Also separate a brand mention from a link to the brand's domain. One does not prove the other.
Avoid a single composite score unless the weighting is documented and useful for a specific decision. Composite scores can hide the exact issue the team needs to fix.
Check Citations and Technical Eligibility

Open every citation in a manageable sample and record three results: whether the URL resolves, whether the page supports the claim, and whether the information is current. Mark uncertain cases for a second reviewer. A working link is not proof that the answer represented the source correctly.
Maintain a short fact sheet for information that assistants often get wrong, such as current product names, pricing page location, supported regions, contact details, and policy dates. Score only facts that can be checked against a named source. Do not fill gaps with assumptions.
Technical checks should follow each platform's published requirements. Google says pages must be indexed and eligible to show a snippet to appear as supporting links in its AI search experiences. Its guidance does not require special AI markup.
OpenAI's publisher guidance explains how OAI-SearchBot controls discovery for ChatGPT search separately from GPTBot training controls.
Check robots directives, indexability, canonical URLs, server responses, visible page content, and internal links. Use structured data only when it matches content readers can see and follows the relevant search documentation. These checks improve eligibility and clarity.
They do not guarantee inclusion or citation.
Choose Tools by Method, Not by Score

Tool coverage and product behavior change quickly. Evaluate the current product rather than relying on an old comparison table. Ask each vendor to show:
- the engines, markets, languages, and product modes covered;
- how prompts are scheduled and whether repeated runs are supported;
- the raw answer and citation evidence behind each score;
- how failed runs, missing citations, and duplicate sources are handled;
- whether results can be exported with timestamps and prompt versions;
- what changes when an engine updates its model or interface;
- the current retention, privacy, and billing terms.
Run a small acceptance test before committing. Give each tool the same prompt set and manually verify a sample of its stored answers and links. If the platform exposes only a score, with no evidence behind it, it is difficult to audit and weak for decision making.
A spreadsheet is enough for an initial baseline. Automation becomes useful when it preserves the raw evidence, repeats the method consistently, and reduces collection work without hiding failures. Budget analyst time for citation and factual review even when collection is automated.
Do the assistants your buyers ask name you, or a competitor?
Reads your site, then asks four assistants what your customers ask.
Interpret Engine Variance Carefully
Different assistants can return different answers because their retrieval systems, indexes, interfaces, models, personalization, and search availability differ. Treat that variance as a reason to inspect evidence, not as proof that one engine permanently prefers a source type.
Compare like with like. Use the same prompt wording, market, language, period, and signed-in state where possible. Report each engine separately before showing an aggregate. If one product returns no citations for a response, record that outcome instead of guessing which source influenced it.
Industry benchmarks can help generate questions, but they are not a substitute for your own baseline. Before using a benchmark, check its prompt set, collection dates, sample selection, geography, engine configuration, and definition of a mention or citation. Scores from different vendors are not directly comparable when the methods differ.
When a result moves, review the underlying answers and cited URLs. A higher mention rate can still include inaccurate claims. A lower citation rate can come from an interface change, a failed collection run, or a genuine source shift. The raw evidence tells you which explanation is plausible.
Connect Visibility to Decisions Without Claiming Causation

A mention is useful brand context, but it does not prove authority, traffic, or revenue. A citation is stronger evidence that a page was presented as a source, yet it still does not prove that a person clicked or converted.
Use analytics to record referral sessions that expose a recognizable AI source. Attach the same lead, signup, or purchase events used for other channels. Document the attribution model and window.
Some users will return through direct traffic or another channel, and some assistants may not pass a useful referrer, so reported AI referrals are an observable subset rather than a complete count.
For changes to content or technical setup, record the intervention date and keep the prompt set stable. Compare pre-change and post-change observations, then check for product updates, seasonality, campaigns, and other explanations. Describe a correlation as a correlation.
A controlled experiment may support a stronger conclusion, but a dashboard trend alone does not.
The report should end with actions tied to evidence. Fix broken pages, correct unsupported claims, update stale facts, improve pages that already earn relevant citations, and investigate prompt groups with repeated factual errors. Do not promise a citation increase.
Measure the next sample and show what changed.
Find out what ChatGPT says about you before your next buyer does.
Free, no account. The paid audit is $490 and takes 3-5 business days.

