Melbourne, Australia 0428 625 394 tasman.murray@holisticanalytics.com.au

Home  /  Insights  /  AI & measurement

How to measure your brand's visibility inside AI answers

Your buyers are building shortlists inside AI assistants before they reach a search results page. Here is a method for measuring whether you appear in those answers, built the way we would build any measurement system.

Tasman Murray  ·  Managing Partner  ·  30 July 2026

Brand visibility inside AI answers is measured by running a fixed set of buyer-representative prompts across each platform on a repeating schedule, and recording three things for every response: whether your brand appears, how it is described, and which sources the model drew on. Because these systems are non-deterministic, a single answer tells you nothing. The unit of measurement is the distribution across many repeated runs, reported separately for each platform.

Why the old measure stopped being enough

Share of Search worked because search volume was a clean, external, weekly-updating proxy for brand demand. Its power came from a simple property: every person researching your category left a countable trace in a keyword tool.

That property is eroding. When someone asks an assistant who the best marketing mix modelling consultancies in Australia are, they run one conversation instead of ten searches and read one synthesised answer naming three or four firms. No keyword tool sees the question. If your brand wasn't in that answer, you were excluded from the shortlist and you have no record that it happened.

Search was a position question: where do we rank. AI answers are a presence question: were we in the response at all, and what did it say about us. Ranking tools were never built to measure presence.

The shift in one line

The five metrics that matter

The industry has settled loosely on Share of Model for the headline number, though you will also see share of citation, AI share of voice and share of answer used for the same idea. A single blended figure is a boardroom number, not a working one. We report five.

1. Presence rate

The proportion of runs mentioning your brand at all, per platform, per prompt category. The base measure, and the first to move when something changes.

2. Recommendation rate

The proportion of runs where you are actively recommended rather than merely mentioned. Being listed among twelve vendors is not the same as being one of three answers to who should I use. The gap between presence and recommendation is usually the most commercially interesting number on the page.

3. Share of Model

Your presence rate as a proportion of the total across your named competitive set. Share is relative by definition, so the competitor list must be defined deliberately and held constant.

4. Characterisation

What the model says about you when it does mention you. This is the metric with the most downside risk. A model confidently describing your pricing, service range or geography incorrectly is doing measurable commercial damage in a channel with no impressions report to alert you.

5. Source attribution

Which properties the model cited or drew on. The actionable metric, because it tells you where to work.

Designing the prompt set

The prompt set is the study. Everything downstream inherits its quality, and this is where in-house attempts usually go wrong. Teams write the prompts they wish buyers asked rather than the ones buyers ask.

Build it from evidence you already hold: sales call notes and CRM free text, existing search query data rewritten into conversational form, and objection language from win/loss records. Then stratify by buying stage.

StratumExample shapeWhat presence here means
Unbranded categoryHow do I work out what my TV advertising is worth?Whether you exist in the model's map of the category
Vendor selectionWho are the best analytics consultancies in Australia?Shortlist inclusion, the highest value stratum
ComparisonFirm A vs Firm B for marketing mix modellingHow you are positioned against named rivals
BrandedWhat does this firm actually do?Accuracy risk. Errors here are urgent
Problem-ledOur attribution broke after signal loss, what now?Whether your thought leadership reaches the model

Fifteen to twenty prompts establishes a baseline. Fifty or more is where a B2B trend line becomes stable enough to act on. Once set, freeze the list. Every prompt you add or reword breaks comparability with everything measured before it.

Sampling: the part most tools skip

Large language models are non-deterministic. Ask the same question five times and you can get five different answers. This invalidates most casual measurement, including the common executive habit of asking once, not seeing the brand, and declaring a crisis.

  • Repeat each prompt across multiple runs per period. Five to ten per prompt per platform is a workable floor. Report the proportion with a confidence interval.
  • Control the session context. Personalisation, memory and chat history all influence output. Run from clean sessions or via API, and document which.
  • Control geography. An answer served from a Melbourne IP differs from one served from London. For an Australian business this is not a detail.
  • Fix the cadence. Weekly or fortnightly. Ad-hoc measurement produces noise you will misread as signal.

A caution on tooling. There is now a crowded market of AI visibility platforms, several genuinely useful for automating this at scale. But the models do not publish what users actually ask or how often, so every AI search volume figure you are shown is an estimate built on assumptions the vendor should be able to explain. Ask them to. If the methodology isn't disclosed, treat the number as directional at best.

What actually moves the number

Here is the finding that reorders most marketing plans: the majority of what a model cites when answering a category question is not your website. It is third-party coverage: industry publications, comparison sites, forums, review platforms and analyst commentary.

Traditional SEO was primarily a first-party game where you optimised your own site. This is primarily a third-party game. That has an uncomfortable organisational implication: the lever with most influence over your Share of Model sits with PR, partnerships and content distribution, not with whoever owns your CMS.

Alongside that, a smaller set of on-site factors reliably help and are cheap: lead with a direct 40 to 60 word answer under a question-shaped heading, write short self-contained blocks rather than flowing narrative, cite sources and include figures, and keep your entity facts consistent across every profile a crawler can reach.

What does not work, currently, is paying for it. No major platform offers paid placement inside generative answers. This is an earned channel, which is why it rewards organisations already investing in credibility.

Connecting it to commercial outcomes

Presence is a leading indicator, not a result. Two honest ways to connect it to the business:

Model it. If you already run marketing mix modelling, Share of Model can enter as a variable once you have enough periods of history. Expect twelve months of consistent measurement before the coefficient means much.

Ask. Add one question to your enquiry form and to sales discovery: how did you first come across us, and did you use an AI assistant while researching. Crude, self-reported, and still the fastest way to size the channel. Most organisations skip it because it feels unsophisticated, then spend six months building attribution for something they could have sized in a fortnight.

Frequently asked questions

What is Share of Model?

Share of Model is the proportion of AI-generated answers mentioning your brand relative to a defined competitor set, across a fixed prompt set. It is the AI-answer equivalent of share of voice, and must be reported per platform because visibility in one model does not transfer to another.

How is it different from Share of Search?

Share of Search measures your proportion of category search volume using observable keyword data. Share of Model measures your proportion of AI answer mentions using data you generate yourself by prompting. One is observed, the other sampled. Related questions, not substitutes.

How often should we measure it?

Weekly or fortnightly on a frozen prompt set with multiple runs per prompt. Citations shift as models retrain and competitors publish, so a quarterly snapshot misses both the erosion and its cause.

Do we need a paid platform to do this?

Not to start. A twenty-prompt baseline run manually across two platforms takes about a day and establishes whether you have a problem worth solving. Paid tooling earns its place when you need frequency, scale and history.

Which platforms should we measure?

The ones your buyers use, which is rarely all of them. For most Australian B2B organisations that means ChatGPT and Google's AI surfaces first, then Perplexity and Gemini. Measure each separately and never report a blended figure without the per-platform breakdown underneath it.

Want this applied to your business?

Tell us the decision you are trying to make. If we are not the right people, we will say so and point you somewhere better.

Start a conversation

More insights