Agencies need to show progress, but AI visibility cannot currently be measured with the same apparent stability as a conventional ranking. Answers vary, source sets change and commercial attribution is incomplete.
The solution is not to avoid measurement. It is to use a balanced scorecard, document limitations and combine several forms of evidence.
The Measurement Problem
The same prompt can produce different answers by platform, location, user and date. Some systems browse the web while others rely more heavily on existing model knowledge. A tracked prompt set is therefore a sample.
Machine Confidence Score
A composite score can summarise technical foundations, entity clarity, evidence, authority, reputation and trust. The inputs and weightings should be documented and version controlled.
The pillar profile is often more useful than the overall number because it reveals the limiting capability.
Prompt Visibility
Track a controlled set of prompts aligned to audiences, buying journeys and commercial themes. Record whether the brand is absent, mentioned, listed, recommended or cited.
Repeated sampling can indicate stability, although it cannot remove volatility.
Share of Answers
Share of Answers can measure the proportion of tracked responses in which the brand appears. More advanced classifications can distinguish positive recommendation, neutral inclusion and negative context.
A mention should not be counted as equal to a recommendation.
Citation Metrics
Track citation frequency, source type, topical relevance, freshness and whether the source supports the client’s claim. Compare the client with competitors and monitor changes in recurring source patterns.
Entity and Evidence Metrics
Measure profile consistency, description accuracy, expert coverage, case-study depth, research assets, verified claims and evidence freshness. These metrics show capability development before it appears in prompt visibility.
Reputation Metrics
Measure review freshness, rating, theme movement, response rate, advocacy assets and community sentiment. Reputation reporting should surface operational issues rather than merely celebrate positive reviews.
Commercial Metrics
Where possible, connect AI referrals, branded demand, leads, pipeline, sales-call mentions and self-reported attribution. The agency should acknowledge that AI influence may appear through direct or conventional search journeys.
Commercial interpretation should combine analytics, CRM and qualitative evidence.
The Monthly Report
The report should answer four questions: what did we do, what changed, what did we learn and what happens next? Outputs should be connected to strategic progress.
Quarterly reviews should reassess priorities rather than repeat the monthly dashboard.