← The journalMEASUREMENT / FIELD GUIDE

Measure AI visibility without fooling yourself

Define mentions, citations, position, and share of voice with denominators that make your reports comparable.

Stockholm street and car with directional camera-motion blur.MEASUREMENT / SUNDIAL
Photography: Danial / Unsplash ↗ · Sundial motion treatment
THE PRACTICAL TAKEAWAY

Every percentage needs a stated denominator and a stable sampling method.

Write the measurement contract first

Before building a dashboard, decide what counts. Does a mention include a product name without the parent brand? Are repeated mentions counted once per answer or multiple times? Does position apply only to explicit ranked lists? How are failed requests treated? These choices change results even when the answers do not change.

A practical contract can count visibility as the share of successfully collected answers containing an eligible brand mention. Report collection failures separately. If you exclude them silently, a provider outage can change the apparent trend by changing which questions were successfully observed.

Keep four events distinct

A mention names a brand. A citation links to a source. A referral brings a visitor to a site. A conversion records a business event. An answer can mention you while citing an independent reviewer; it can cite your documentation without recommending your product. Neither event proves a customer visited or bought.

Define share of voice explicitly. One valid measure is your brand’s qualifying mentions divided by all qualifying mentions among a fixed competitor set. That differs from answer-level visibility. Label both precisely and retain the competitor list used in each period.

Use an illustrative calculation

Suppose a fictional test collects 40 successful answers from 50 scheduled attempts. Your brand appears in 12 of the successful answers. Under the answer-level definition above, visibility is 30%, with an 80% collection success rate. It is not 24% unless your stated method deliberately includes unsuccessful attempts in its denominator.

If eight of those 12 answers contain an ordered shortlist, calculate average rank on the eight eligible lists and disclose that sample size. Assigning an arbitrary last-place rank to the four unranked answers creates precision the source material never provided.

Make comparisons reproducible

Save exact prompts, timestamps, model or product identifiers where available, markets, languages, response text, and citation URLs. Keep the baseline prompt set fixed. Explain additions and removals. Compare like-for-like slices before attributing a movement to your campaign.

Show sample sizes next to the chart and link to representative answers. A small change based on a handful of observations deserves inspection, not a success headline. Your report should help a colleague reproduce the conclusion without trusting the visualization alone.