Brand visibility used to be a ranking question. You checked position three, position seven, position eleven, and you knew where you stood. That instrument no longer covers the whole market, because a growing share of commercial questions are now answered inside ChatGPT, Gemini and Perplexity — as prose, with two or three brands named in it and nobody else mentioned at all.

The uncomfortable part is that most companies have no idea whether they are one of the named brands. They are not in the answer, they are not absent from it, they are simply unmeasured. When a board asks “are we showing up in AI?”, the honest answer in most organisations is a shrug, and a shrug is not a strategy. This guide replaces the shrug with a method: a fixed prompt set, a scoring rubric, and a number you can put on a slide and defend.

None of this requires exotic tooling. A large language model answering a buying question is doing something you can observe directly, and a spreadsheet plus two hours a month will get you a defensible baseline. Paid platforms make it faster and they do not make it truer. What follows is the manual method first, because if you cannot describe how your brand visibility number is produced, you cannot tell whether a tool is producing it correctly.

Everything here is method, not magic. Each chart below is arithmetic performed on figures stated in the surrounding text, so you can substitute your own counts and rerun it. If you want the strategy layer that sits on top of the measurement, our GEO services and AEO services pages cover how the optimisation work is structured once you can see the score.

Why Brand Visibility in AI Answers Needs Its Own Measurement

measure brand visibility chatgpt gemini perplexity b three identical upright cylinders

The instinct is to bolt this onto the existing SEO report. That instinct is wrong, and understanding exactly why is what stops you from measuring the wrong thing for six months.

There is no rank, so there is no rank tracker

A generative answer has no positions. There is no slot two that you can occupy next quarter. There is a paragraph, and your brand is either in it or it is not. Every metric you are used to — average position, movement, pixel depth — has no referent here. The unit of brand visibility is the mention, not the placement.

The same question produces different answers

Ask ChatGPT the same commercial question three times and you can get three materially different answers, naming different companies. This is not a bug you can engineer around; it is how sampling works. It has one enormous consequence for measurement: a single check tells you almost nothing. Any brand visibility figure built on one run of one prompt is noise dressed as data.

The traffic arrives with the question stripped off

When a user reads about you in Gemini and then visits your site, the query that produced the mention does not travel with them. Referrer data is thin, often absent, and increasingly consolidated. You cannot reverse-engineer brand visibility from analytics, which is precisely why you have to observe the answers directly rather than infer them from traffic.

The competitive set is smaller and harder

Ten organic results is a generous market. Three named brands in a paragraph is a brutal one. Coming fourth in a generative answer is identical to coming fortieth — you are not in the text. That compression is the single strongest argument for measuring brand visibility separately and early, before a competitor’s position calcifies into the model’s default answer.

Classic SEO metricAI answer equivalentWhat changes
Average positionPresence rateBinary in/out per answer, averaged over many runs
ImpressionsPrompt coverageYou define the demand set; nobody reports it to you
Click-through rateCitation rateA linked source, not a blue link the user chose
Competitor rank gapShare of voiceYour mentions as a share of all brand mentions
SERP feature ownershipFraming and sentimentHow you are described, not merely whether
Rank volatilityAnswer varianceRun-to-run instability is expected, not an anomaly

Read that table as a translation guide rather than a replacement plan. The classic metrics still run your organic programme; the right-hand column runs a second, parallel programme measuring brand visibility inside assistants.

What Brand Visibility Means Inside ChatGPT, Gemini and Perplexity

measure brand visibility chatgpt gemini perplexity c tall stack blank paper sheets

“Are we visible?” is four different questions wearing one coat. Separating them is the first real work, because they have different causes and different fixes.

A mention is your name appearing in the prose

The model wrote your brand name in the answer body. No link required. This is the base layer of brand visibility and the one most people mean when they ask the question — it is also the one that a purely link-based tracking approach will miss entirely.

A citation is your domain appearing as a source

The answer carries a footnote, a source chip or an inline link pointing at your site. Perplexity does this most visibly, ChatGPT does it when browsing, and Gemini surfaces it through its own link treatments. A citation is worth more than a mention because it is the only one that can send you a visitor.

A recommendation is being named as the answer

There is a wide gap between “vendors in this space include A, B and C” and “for a business of that size, B is usually the right starting point”. The second is a recommendation. Tracking brand visibility without separating these two collapses your best and worst outcomes into the same tally.

Framing is what the sentence around your name says

Being named as “the budget option”, “the enterprise-heavy one” or “the specialist for regulated industries” shapes the buyer before they ever reach you. Framing is the part of brand visibility that most resembles PR, and it is the part that a numeric score alone will hide from you.

StateWhat you see in the answerCommercial valueHow to record it
AbsentNo trace of your brandNoneScore 0
MentionedName in the prose, no linkAwareness, no clickScore 1
CitedYour domain listed as a sourceAwareness plus referralScore 2
RecommendedNamed as the preferred fitHighest — pre-qualified intentScore 3
MisframedNamed with a wrong or dated claimNegativeFlag separately, do not score

Keep the misframed flag out of the numeric score. A factual error about your pricing or your service range needs a correction workflow, not a percentage point, and burying it inside a brand visibility average is how it goes unfixed for a year.

The Five Brand Visibility Metrics Worth Tracking

measure brand visibility chatgpt gemini perplexity d upright funnel

Five numbers cover the ground. More than five and nobody reads the report; fewer and you cannot tell a coverage problem from a preference problem.

Presence rate

The share of answers in which your brand appears at all. If your prompt set is 40 prompts run three times on one engine, that is 120 answers; appearing in 30 of them is a presence rate of 25%. This is the headline brand visibility number and the one to trend month over month.

Share of voice

Your mentions as a proportion of all brand mentions across the same answer set. Presence rate tells you whether you are in the room; share of voice tells you how crowded the room is. A rising presence rate with a flat share of voice means the category is getting more coverage and you are merely keeping up.

Citation rate

The share of answers that link to your domain. Track it separately from mentions, because the two move for different reasons: mentions respond to how well known you are, citations respond to whether your pages are retrievable, quotable and current.

Answer position

Where in the answer your brand first appears — first named, mid-list, or a trailing “others include”. First-named is disproportionately valuable, and a brand visibility report that treats all mentions equally will show a flat line while your position quietly degrades.

Framing score

A simple three-way tag on each mention: positive, neutral or inaccurate. It resists automation and it is worth the manual effort, because it is the only metric that catches a model confidently describing a service you stopped selling two years ago.

MetricHow it is calculatedReported asMoves when
Presence rateAnswers containing you ÷ total answersPercentageCategory authority grows or decays
Share of voiceYour mentions ÷ all brand mentionsPercentageCompetitors gain or lose ground
Citation rateAnswers linking your domain ÷ total answersPercentagePages become retrievable or stale
Answer positionMean rank of first appearanceOrdinal averagePreference shifts within the set
Framing scorePositive and neutral tags ÷ all mentionsPercentage plus an error logPublic source material changes

Report all five together. Any one of them read alone can be moved by something that has nothing to do with your brand visibility work, and the pattern across the set is what tells you which.

Building the Prompt Set Your Brand Visibility Score Depends On

measure brand visibility chatgpt gemini perplexity e three hexagonal slabs

The prompt set is the measurement instrument. Get it wrong and every number downstream is wrong in the same direction, permanently, and no amount of scoring rigour will rescue it.

Start from the buying journey, not from your keyword list

Keyword research produces fragments — “crm software uk”, “managed it chester”. People do not type fragments into an assistant; they type sentences with context in them. Your prompt set should read like the questions a buyer actually asks a knowledgeable friend, because that is the register these tools were trained to answer.

Use four prompt categories

Split the set evenly. Category prompts (“what tools do mid-sized UK manufacturers use for X”) test whether you exist in the category at all. Comparison prompts (“X versus Y for a 200-person firm”) test preference. Problem prompts (“our stock system keeps drifting out of sync, what do we do”) test whether you surface without the category being named. Brand prompts (“what does [your brand] do”) test accuracy rather than reach.

Forty prompts is the working number

Ten per category gives you enough resolution to see movement without turning collection into a project. At three runs per prompt on three engines, 40 prompts produce 360 answers per cycle — large enough that a single odd response cannot swing your brand visibility figure, small enough to finish in an afternoon.

Freeze the wording and never edit it

The moment you reword a prompt, its history is gone. Treat the set as a locked instrument: if a prompt genuinely needs replacing, add the replacement as a new line and retire the old one with a date, exactly as you would version a survey question.

Record the conditions with every run

Log the engine, the model version if it is shown, the date, whether search or browsing was enabled, and the account state. These conditions change results more than anything you will do to your website that month, and a brand visibility trend that ignores them is comparing incomparable things.

Worked example: presence rate by prompt category, 10 prompts per category
Brand prompts — 7 of 10 70%
Category prompts — 4 of 10 40%
Comparison prompts — 2 of 10 20%
Problem prompts — 1 of 10 10%

That shape — strong on brand prompts, weak on problem prompts — is the most common brand visibility profile there is, and it has a precise meaning. The model knows who you are once your name is supplied, and does not associate you with the problems you solve. That is a content gap, not an authority gap, and it is comparatively cheap to close.

Running the Test: A Repeatable Monthly Brand Visibility Method

measure brand visibility chatgpt gemini perplexity f upright hourglass

Consistency beats sophistication here. A crude method run identically every month produces a usable trend; a sophisticated method run differently each time produces nothing.

Use a clean, logged-out session

Personalisation and memory will quietly show you a flattering answer. Use a fresh private window, a logged-out state or a dedicated measurement account with memory and history disabled, and note which you chose. Mixing states across months is the single most common way a brand visibility trend becomes fiction.

Run every prompt three times

Three runs let you record a fraction rather than a coin flip: appearing in two of three is meaningfully different from one of three. It also surfaces the variance itself, which is worth reporting — a brand that appears erratically is in a different position from one that appears reliably.

Capture full text, not screenshots

Paste the complete answer into a cell. Text is searchable, diffable and countable; a screenshot is none of those. Six months in, the archive of raw answers is more valuable than any chart you drew from it, because it lets you re-score against a rubric you had not thought of yet.

Log citations as a separate list

Every source the answer displays, whether it points at you or not, goes into its own column. The competitor citation list is the most actionable output of the whole exercise: it is a literal inventory of the pages the model trusts on your topic.

Budget the time honestly

At 40 prompts, three runs and three engines, one cycle is 360 answers. At roughly 45 seconds per answer to run, paste and tag, that is 270 minutes — about four and a half hours a month. Halving the runs to one halves the confidence and saves 180 minutes. Adding a 30-second citation log to every answer adds another 180 minutes.

Minutes per monthly cycle, from the figures stated above (40 prompts, 3 engines)
Three runs plus citation logging — 270 + 180 450 min
Three runs, scoring only — 360 answers at 45s 270 min
One run, scoring only — 120 answers at 45s 90 min

Four and a half hours a month is less than one working day, and it is the whole cost of having a defensible brand visibility number instead of an opinion. Most teams that abandon the exercise do so because they tried to measure 200 prompts in month one.

Scoring: Turning Raw Answers Into a Brand Visibility Number

Raw answers are evidence. The score is what makes them comparable across months, engines and competitors, and the rubric matters more than the arithmetic.

Use the 0 to 3 rubric

Absent scores 0, mentioned scores 1, cited scores 2, recommended scores 3. Sum across the answer set and divide by the maximum possible. Forty prompts scoring a theoretical maximum of 3 each gives 120 points per run; a total of 34 is a brand visibility index of 28%.

Weight by prompt category if you must, but declare it

If comparison prompts matter more to your pipeline than brand prompts, weight them — and write the weighting into the method document. An unweighted score that everyone understands beats a weighted one that only its author can reproduce.

Keep competitor scores in the same sheet

Score the top three competitors on the same rubric from the same answers. It costs almost nothing extra because you are already reading the text, and it converts an abstract percentage into a ranking your leadership will actually act on.

Worked share-of-voice example

Suppose one monthly cycle produces 57 brand mentions in total across the answer set. Your brand accounts for 12 of them, competitor A for 18, competitor B for 15, competitor C for 8, and four mentions are spread across everyone else. Dividing each by 57 gives the share-of-voice split below.

Worked example: share of voice from 57 total brand mentions
Competitor A — 18 of 57 32%
Competitor B — 15 of 57 26%
Your brand — 12 of 57 21%
Competitor C — 8 of 57 14%
Everyone else — 4 of 57 7%

Third place on 21% with the leader on 32% is a recoverable position; the same 21% against a leader on 60% is a different strategic problem entirely. That is why share of voice belongs next to presence rate in every brand visibility report rather than in a separate appendix.

How the Three Engines Differ, and Why One Score Is Not Enough

Averaging the three engines into a single figure destroys the most useful signal in the data, which is the disagreement between them.

ChatGPT

It answers from trained knowledge unless it decides to browse, so brand visibility here rewards long-standing, widely-replicated presence across the open web. Its search behaviour means a prompt phrased as a current-events question is far more likely to fetch live pages than one phrased as general advice — so record which mode produced each answer.

Gemini

Tied closely to Google’s index and surfaced alongside AI features in Search itself. In practice, brand visibility in Gemini correlates more with conventional organic strength than the other two do, which makes it the engine where existing SEO investment shows up first. The technical SEO fundamentals that drive organic ranking are doing double duty here.

Perplexity

Built around retrieval and citation, so it shows its sources by default and updates fastest. It is the best early-warning system of the three: a new page that earns citations in Perplexity within days is usually the same page that starts appearing in the other two engines weeks later.

FactorChatGPTGeminiPerplexity
Primary source of answersTrained knowledge, browsing when triggeredGoogle index plus model knowledgeLive retrieval first
Citations shownWhen browsingVia link treatmentsBy default, prominently
Speed of reflecting new contentSlowestMedium, follows indexingFastest
What lifts brand visibility mostBroad third-party corroborationConventional organic strengthQuotable, current, crawlable pages
Best used in your report asLagging authority indicatorBridge to your SEO programmeLeading indicator of change

Report the three side by side and read the gaps. Strong Perplexity numbers with weak ChatGPT numbers usually mean recent work that has not propagated yet — which is encouraging. The reverse pattern means historical reputation carrying you while your current pages go uncited, which is a warning.

Manual Tracking, Platforms or Log Analysis for Brand Visibility

Three approaches exist and they answer different questions. Most teams should start with the first and add the third, and consider the second only once they know what they want from it.

The manual spreadsheet

One sheet, one row per answer, columns for engine, prompt, run number, date, score, competitors named and citations listed. It is slow, it is completely transparent, and it produces the raw archive that makes every later analysis possible. Start here even if you intend to buy a platform.

Dedicated AI visibility platforms

A crowded and fast-moving category of tools that automate the running and scoring of prompt sets. They save real time and their headline numbers are only as good as their prompt sets and their sampling frequency, both of which you should interrogate before signing. Our overview of one such GEO visibility platform walks through what these products actually do.

Server logs and analytics

The only approach that measures outcomes rather than appearances. It tells you nothing about answers you were absent from, which is exactly the gap the other two methods fill — but it is the only one that connects brand visibility to revenue.

ApproachMonthly effortAnswers the questionMain weakness
Manual spreadsheet4–8 hoursAre we named, cited or recommended?Labour, and a ceiling on prompt volume
Visibility platform1–2 hours plus licenceHow is the trend moving at scale?Opaque prompt sets and sampling
Logs and analytics1–2 hoursDid any of it produce visitors?Silent about answers you never appeared in
All three combined6–10 hoursAppearance, trend and outcomeNeeds one owner or it fragments

If you only ever do two of the three, do the spreadsheet and the logs. Between them you get the appearance side and the outcome side of brand visibility, and a platform then becomes a scaling decision rather than a measurement decision.

Reading Logs and Analytics for AI Brand Visibility Signals

Your own server is an underused instrument here, and unlike the answers themselves it produces data nobody has to interpret subjectively.

Segment assistant referrals in analytics

Create a channel grouping for referrals from the assistant domains and watch it as its own line. The volume will look small next to organic for a long time; the conversion quality frequently does not, because a visitor arriving from a recommendation has already been pre-qualified by the answer.

Watch the AI crawlers in your access logs

Requests from the documented crawler user agents tell you which of your pages are being fetched for training and for live retrieval. Google publishes its common crawlers list and OpenAI publishes its bot documentation, so the identification is straightforward. A page nothing fetches cannot contribute to brand visibility.

Check what you are blocking before you diagnose anything else

More than one company has spent a quarter puzzling over poor brand visibility while a robots directive quietly excluded the retrieval bots. Read your own file against the robots exclusion standard first — it is a five-minute check that occasionally explains everything.

Accept that attribution will be incomplete

Some assistant traffic arrives with no referrer at all and lands in direct. Rather than fight it, watch direct traffic to deep pages that nobody bookmarks and nobody types — a sustained rise there, correlated with rising presence rate, is about as close to proof as this channel currently offers.

What Actually Moves Brand Visibility Once You Can Measure It

Measurement without levers is just anxiety with a spreadsheet. These are the five things that reliably change the numbers, in rough order of effort.

Be quotable in the first place

Models reproduce sentences that state something specific and self-contained. A page that answers a question in one clear paragraph near the top gets used; a page that warms up for 400 words does not. This is the cheapest brand visibility improvement available and it is mostly an editing job.

Make the entity unambiguous

Consistent naming, a clear description of what you do, structured data via organisation markup, and matching details everywhere you appear. Models resolve entities from corroboration across sources; contradictory descriptions dilute the association and cap your brand visibility regardless of content quality.

Earn third-party corroboration

Being described accurately on sites you do not control is worth more than anything on your own domain, because it is what turns a claim into a consensus. Directory listings, industry roundups, comparison sites, credible press and genuine reviews all feed the same mechanism.

Keep the important pages current

Retrieval-based engines strongly prefer recent, dated material. A page updated this quarter has a structural advantage over a better page from 2023, and a documented content maintenance routine is what converts that into a durable brand visibility gain rather than a one-off spike.

Keep the technical door open

Crawlable, fast, server-rendered pages with clean markup. If a retrieval bot cannot fetch and parse the page, none of the above matters — and the Core Web Vitals and rendering work you may already have in flight for organic search serves this directly.

Brand Visibility Measurement Mistakes That Waste a Quarter

Every one of these has been made by a competent team. They are systematic, not careless.

Measuring once and declaring a result

A single run against a single prompt is a sample of one from a probabilistic system. It is genuinely worse than no measurement, because it produces a confident number that will not reproduce.

Writing prompts that beg for your name

“What are the best IT support companies in Chester like Progressive Robot?” is not a measurement, it is a hint. If your prompt contains your brand name, it belongs in the brand-accuracy category and must never be counted towards category brand visibility.

Measuring in a personalised account

Your own logged-in session knows your site, your history and possibly your previous questions about yourself. Every measurement it produces is inflated, and the inflation is invisible in the output.

Reporting a single blended number

One figure across three engines hides the disagreement that carries all the diagnostic value. Report per engine, always, and reserve the blended number for the summary slide.

Treating it as a replacement for SEO

Almost everything that improves brand visibility in these tools also improves conventional organic performance, and much of it is conventional organic work. Treating this as a separate discipline with a separate budget and a separate team is how organisations end up doing the same job twice. Our SEO services and AIO services sit in one programme for exactly this reason.

A 90-Day Brand Visibility Measurement Plan

Ninety days is enough to establish a baseline, take one action and observe whether it moved anything. Less than that and you are reading noise.

Days 1 to 14: build and freeze the instrument

Write the 40 prompts across the four categories. Name your three competitors. Build the sheet. Run the first full cycle across all three engines and record the conditions. This first cycle is your baseline and it is the only one whose numbers mean nothing on their own.

Days 15 to 45: act on the citation inventory

Read the competitor citation list you collected. It names the exact pages the models trust on your topic. Build or improve the two or three pages where you have the strongest genuine claim to be the better source, and fix any factual errors the answers revealed about you.

Days 46 to 60: run cycle two and compare

Same prompts, same conditions, same rubric. Compare per engine. Expect Perplexity to move first and ChatGPT to be unchanged — that is the normal shape of a brand visibility response, not a failure.

Days 61 to 90: run cycle three and decide the cadence

With three cycles you have a trend rather than two points. At this stage decide whether monthly manual measurement is sustainable, whether a platform is worth its licence, and which single metric goes on the leadership dashboard. For most B2B firms with a long buying cycle, presence rate per engine plus share of voice is the right pair, and our B2B search strategy guide covers how that fits a longer pipeline.

Frequently Asked Questions About Brand Visibility in AI Answers

How often should we measure?

Monthly for most organisations. Weekly only if you are running an active campaign and need to catch movement quickly; quarterly is too infrequent to separate a real change from ordinary answer variance.

Can we pay to appear in AI answers?

Not in the organic answer text itself in any reliable way. Advertising products are appearing alongside generative results, but the brand visibility discussed in this guide is earned through the source material the models draw on, which is why the levers look like content and PR rather than media buying.

How long before improvements show up?

Perplexity and other retrieval-first tools can reflect a new page within days. Gemini tends to follow indexing. ChatGPT’s trained knowledge moves slowest of the three, which is why a realistic brand visibility programme is measured over quarters, much like the timeline for conventional SEO results.

Is a low score a content problem or an authority problem?

Look at the category split. Weak on brand prompts is an accuracy problem. Weak on problem prompts with strong brand prompts is a content gap. Weak everywhere, including brand prompts, is an entity or authority problem and takes longest to fix.

Should small businesses bother?

Yes, and often more than large ones. In a narrow local or specialist category the number of credible brands is small, so the effort required to become one of the two or three names an assistant returns is far lower than competing for the same position in a national market.

References