Skip to main content Scroll Top

LLM Visibility: What It Measures, and Why Most Numbers Are Not Reproducible

Updated September 2026 · Written and maintained by the Progression Agency strategy team

Four separate measurements that get collapsed into one misleading score — and the measurement record that makes an AI visibility claim falsifiable instead of assertable.

On this page · 12 sections
  1. What LLM visibility means
  2. Why most reported LLM visibility numbers are not reproducible
  3. What actually moves the number
  4. How to build a prompt set that survives scrutiny
  5. What to expect and when
  6. Claims that should not be believed
  7. Why the LLM-prefixed terms matter commercially
  8. Everything else we have written on search, AI and getting found
  9. The four measurements, and what each one tells you to do
  10. A measurement method you could hand to an auditor
  11. What moves LLM visibility, ranked by effect per pound
  12. LLM visibility versus the terms it gets confused with

The short answerLLM visibility is how often, and how accurately, a brand is named as a source inside AI-generated answers. It is four measurements — presence, share of voice, sentiment and prompt coverage — and merging them into one score makes a decline impossible to diagnose. Because answers vary between runs of the identical prompt, a number is only meaningful if the prompt wording is fixed and versioned, every answer is stored verbatim with platform and date, each prompt is run several times, and the aggregation rule was written down beforehand.

Search volumes, CPC and difficulty measured September 2026. The run-to-run variance chart on this page is illustrative of the structural point, not measured data.

What LLM visibility actually measuresWhat LLM visibility actually measures
Most reported ‘LLM visibility scores’ merge the first four into a single number, which makes a decline impossible to diagnose.

What LLM visibility means

LLM visibility is how often, and how accurately, a brand is named as a source inside answers generated by large language models. It is four separate measurements, not one: presence, share of voice, sentiment and prompt coverage.

The reason to keep them separate is diagnostic. A brand can be present on plenty of prompts but hold a tiny share of voice against a dominant competitor. Another can have excellent share of voice on five questions and appear on nothing else. Those are different problems with different fixes, and a single blended score makes both invisible.

There is no position to measure. An answer names three or four sources and there is no second page, so LLM visibility has no equivalent of ranking fourth — you are named or you are not.

Presence — Metric. Named at all, across the prompt set.
Share of voice — Metric. Your mentions against competitors'.
Sentiment — Metric. Whether the description is accurate and favourable.
Prompt coverage — Metric. Which questions you appear on at all.
Citation vs mention — Metric. Linked versus merely named.
Stability — Metric. How consistently you appear across repeated runs.

Presence

Whether you are named at all, across a fixed set of questions. The blunt first measurement, and the one most sites fail outright.

Share of voice

How often you are named relative to the competitors who also appear. This needs competitor names in some of your prompts, or it has no denominator.

Sentiment and accuracy

When you are named, is the description correct? Being named alongside a wrong claim about your pricing or services is worse than not being named at all.

Prompt coverage

Which buyer questions you appear on and which you are absent from. This is where the actionable gaps live, and it is the measurement most often skipped.

Citation versus mention

A citation carries a link; a mention names you without one. Both have value, they are not the same thing, and they should be counted separately.

Stability

How consistently you appear across repeated runs of the same prompt. A source appearing in one run out of five is not visible in any meaningful sense.

Why most reported LLM visibility numbers are not reproducible

Because answers vary between runs of the identical prompt, and almost nobody stores the answer text. Without it, a drop from twelve citations to eight is indistinguishable from ordinary variance, and nobody — including the supplier reporting it — can tell which happened.

This is the central integrity problem in the discipline right now. It is not usually dishonesty. It is a measurement method that was never designed to be falsifiable.

Is your LLM visibility measurement reproducible?Is your LLM visibility measurement reproducible?
Four reds. Any one of them makes a reported trend unfalsifiable, which is the central integrity problem in this discipline right now.
Why the same prompt gives different answersWhy the same prompt gives different answers
Illustrative, not measured data. The point is structural: a source appearing in three of five runs is a real finding, and a single-run check cannot tell the difference between that and a source appearing once.
Fix the wording — Rule. One changed word changes retrieval.
Store verbatim — Rule. Counts alone cannot be audited.
Record platform and date — Rule. Both change answers materially.
Multiple runs — Rule. Single runs cannot distinguish signal from noise.
State aggregation first — Rule. Decide the method before seeing results.
Never edit a live set — Rule. An edit creates a new set, not an update.

Variance is structural, not a bug

Retrieval is influenced by phrasing, timing, region, account state and model version. Identical inputs genuinely produce different outputs.

One run proves nothing

A source appearing once in five runs and a source appearing in three look identical if you only checked once.

Stored text is the fix

Keep the full answer, dated and platform-labelled. Then a change can be inspected rather than asserted.

Fixed wording is the other fix

Changing one word changes which sources get retrieved. If this month’s prompts differ from last month’s, there is no trend, only two snapshots.

Decide aggregation first

Write down how you will combine multiple runs before you look at the results. Choosing afterwards is how a flat quarter becomes a success story.

What actually moves the number

Crawler access moves it fastest and furthest, costs nothing, and is binary. Rendering moves it a great deal but slowly and expensively. Corroboration moves it most in the long run and cannot be rushed. Publishing more articles barely moves it at all.

The mismatch between where budget usually goes and where the effect actually is remains the most expensive misunderstanding in this field.

What moves LLM visibility, and how fastWhat moves LLM visibility, and how fast
Top-right is where to start. Bottom-right is where most budget goes, which is the mismatch this whole page exists to correct.
Crawler access — Driver. Fast, free, binary. Start here.
Rendering — Driver. Large effect, expensive, slower to show.
Passage structure — Driver. Moderate effect, cheap, weeks not months.
Entity consistency — Driver. Slow, and it is what stops wrong answers.
Corroboration — Driver. Slowest, largest long-run effect.
Content volume — Driver. Minimal effect. Not the lever here.

How to build a prompt set that survives scrutiny

Thirty questions in real buyer language, spread across definition, comparison, provider and problem intents, with competitor names in some of them so share of voice has a denominator. Fix the wording, version it, and record platform, region and date on every run.

None of this is sophisticated. It is ordinary measurement hygiene, and it is absent from most AI visibility reporting we have seen.

How to build a prompt set that survives scrutinyHow to build a prompt set that survives scrutiny
The last rule is the one that separates measurement from storytelling: decide the method before you see the numbers.
Write 30 prompts — Start. Real buyer language, across the funnel.
Pick 2-3 platforms — Start. Named, and held constant.
Run each 3-5 times — Start. In one sitting, same day.
Store every answer — Start. Full text, dated, platform labelled.
Count four things — Start. Presence, share of voice, sentiment, coverage.
Repeat monthly — Start. Identical wording, or it is a new set.

Take the wording from your buyers

Sales call transcripts and support tickets, not a keyword tool. People ask assistants in full sentences and the phrasing matters to retrieval.

Cover four intent types

What is X, X versus Y, who provides X, and why is X not working. Each behaves differently and a set weighted to one of them will mislead you.

Name competitors in some prompts

‘Best X provider’ returns a list. Without competitor names in your set you cannot compute share of voice at all.

Version the set

Give it a number. When you change a prompt you have created version two, and version one’s history does not transfer.

Hold the variables constant

Same platforms, same region, same account state. Each one moves answers enough to manufacture a trend that is not there.

Run each prompt three to five times

In one sitting. Then aggregate by a rule you wrote down beforehand.

What to expect and when

Crawler access shows in your server logs within days. Rendering fixes take a week or two to land. Structure improves extraction over weeks. Presence on niche prompts moves in month two or three. Share of voice on competitive prompts takes three to six months.

Nobody can compress the back half. Corroboration accumulates at the speed third parties publish, and model refresh cycles are outside everyone’s control.

What to expect, and whenWhat to expect, and when
Nobody can compress the back half of this, because corroboration and model refresh cycles are not under anyone’s control.
The numbers that define LLM visibility as a metricThe numbers that define LLM visibility as a metric
The second number is the one that changes everything. With no positions, there is no partial credit and no long tail.

Claims that should not be believed

Four in particular: a guaranteed citation, knowledge of why a model chose a competitor, a single proprietary score that captures true visibility, and a large percentage improvement over a short window.

The last one is the most common and the easiest to check. A 400% improvement over three weeks on a thirty-prompt set is a swing of a handful of mentions. On a metric with this much run-to-run variance, that is noise unless the stored answers show otherwise — so ask to see them.

Claims about LLM visibility that should not be believedClaims about LLM visibility that should not be believed
The 400% claim is the tell. Over three weeks on a small prompt set, a swing of that size is almost always variance or a changed prompt, not progress.
Guaranteed citations — False. Nobody controls model output.
We know the model's reasoning — False. No provider exposes it.
One proprietary score — False. Merging metrics hides diagnosis.
400% in three weeks — False. Almost always variance or a changed prompt.
llms.txt is the lever — False. Low adoption, not a ranking factor.
Rank tracking shows AI visibility — False. Different surface entirely.

On guarantees

Nobody controls model output. A supplier promising a specific citation is describing something that does not exist as a purchasable outcome.

On reasoning

No provider exposes why a model selected one source over another. Every claim to explain it is inference, and inference presented as measurement is the field’s main integrity problem.

On proprietary scores

A merged number is convenient for a slide and useless for diagnosis. It hides which of the four measurements moved and which platform is failing.

On large short-term gains

Ask for the stored answers from both runs. A real improvement survives that request comfortably.

On ‘all major AI engines’

Coverage and reliability vary enormously by platform. Ask for them named individually, with the caveats stated.

Why the LLM-prefixed terms matter commercially

Because they are consistently less contested than their AI-prefixed equivalents while carrying comparable or higher cost per click. ‘LLM visibility tool’ sits at $82.21 a click, the highest figure anywhere in this research, and ‘llm visibility checker’ has a difficulty score of twelve.

That gap will not persist. It exists because the industry standardised its marketing language on ‘AI’ while a technical audience kept using ‘LLM’, and the pages serving the second audience are thinner.

What the market pays for these termsWhat the market pays for these terms
An $82 cost per click on ‘llm visibility tool’ is the highest figure anywhere in this research. These are buyers with budget, not researchers.
How contested each term isHow contested each term is
Difficulty of 12 on a term with a $19 cost per click is effectively an open door. The LLM-prefixed variants are consistently less contested than the ai-prefixed equivalents.
llm visibility tool — Term. 720/mo at $82.21 — highest CPC found.
ai visibility tools — Term. 1,600/mo at $60.46.
ai visibility platform — Term. 590/mo at $56.11, SD 30.
ai visibility optimization — Term. 170/mo at $78.47.
ai visibility audit — Term. 110/mo at $62.08.
llm visibility checker — Term. 90/mo at $19.22, SD 12 — open.

Get a baseline you could defend in a meeting

We build the prompt set in your buyers’ language, run it properly, store every answer verbatim and hand you the whole thing. You keep it whether or not you work with us afterwards.

/ai-visibility-audit

Want this done for your site?We build and maintain the search, content and paid programmes described on this page.

Get a free proposal

Everything else we have written on search, AI and getting found

AI, AEO and what is changing

Websites and design

Choosing and working with an agency

Social, content and brand

By industry and by situation

Not sure which of these applies to you?Tell us the situation and we will say plainly what we would do first, and what we would not.

Talk it through

The four measurements, and what each one tells you to do

Four distinct things people mean by LLM visibility, and the action each one implies.

Keep them separate or you cannot diagnose anything
MeasurementWhat it answersA bad result meansThe fix
PresenceAre we named at all?Access, rendering or retrievability is brokenCrawler access, then rendering
Share of voiceHow often versus competitors?We are retrievable but not preferredPassage structure and corroboration
Sentiment/accuracyIs the description right?Sources disagree about usEntity consistency work
Prompt coverageWhich questions do we appear on?Topical gaps in what we have publishedTargeted content, not volume
Citation vs mentionLinked, or just named?Named without attribution linkUsually nothing — both have value
StabilityConsistent across runs?Marginal retrievabilityStrengthen the passage, not the page count

A measurement method you could hand to an auditor

The full method written out, so that somebody else could reproduce the numbers without asking us anything.

The minimum record for every run
FieldWhy it is requiredWhat goes wrong without it
Prompt text, verbatimWording changes retrievalTrend compares two different questions
Prompt set versionEdits create a new setHistory is silently invalidated
PlatformBehaviour differs materiallyCross-platform averages hide a total failure
RegionAnswers are location-sensitiveResults are not comparable month to month
Date and timeModel versions changeA model update reads as your progress
Run numberVariance is structuralOne run cannot distinguish signal from noise
Full answer textCounts cannot be auditedNo claim afterwards is falsifiable
Aggregation ruleDecided before resultsMethod chosen to flatter the outcome

What moves LLM visibility, ranked by effect per pound

The levers ranked by what they return relative to what they cost, rather than by how often they are sold.

Effort against effect
LeverEffortSpeed to showSize of effectVerdict
Unblock AI crawlersMinutesDaysVery largeAlways first
Fix scriptless renderingHigh1-2 weeksVery largeThe real budget decision
Restructure passagesModerateWeeksModerateReliable, also helps readers
Reconcile schemaLowWeeksSmallCheap hygiene
Unify entity descriptionModerateMonthsModeratePrevents wrong answers
Third-party corroborationHighMonthsLargeSlowest, most durable
Add llms.txtMinutes—NegligibleHygiene, not strategy
Publish more articlesHighMonthsSmallNot the lever here

LLM visibility versus the terms it gets confused with

Four adjacent terms separated, because they are used interchangeably and mean different things.

Five adjacent things that are not the same
TermWhat it isHow it differs from LLM visibility
RankingsPosition among search linksDifferent surface; cannot measure answers
Brand monitoringMentions across web and socialNot restricted to generated answers
Share of searchBrand query volume versus competitorsMeasures demand, not citation
AI referral trafficVisits arriving from AI surfacesAn outcome of visibility, not the measure
AEOThe work done to improve itLLM visibility is the metric; AEO is the practice

AI, AEO and what is changing

Frequently asked questions

What is LLM visibility?
How often, and how accurately, a brand is named as a source inside answers generated by large language models. It is four separate measurements — presence, share of voice, sentiment and prompt coverage — not a single score.
How is LLM visibility measured?
Against a fixed prompt set re-run with identical wording, with every answer stored verbatim and labelled by platform, region and date. There is no position inside a generated answer for a rank tracker to track.
Why do my LLM visibility numbers change every time I check?
Because variance is structural. Retrieval is influenced by phrasing, timing, region, account state and model version, so identical prompts genuinely return different sources. This is why single-run checks prove nothing.
How many runs per prompt do I need?
Three to five, in one sitting, aggregated by a rule you wrote down before seeing the results. A source appearing once in five runs is not visible in any meaningful sense.
How many prompts should be in my set?
About thirty to start, in real buyer language, spread across definition, comparison, provider and problem intents, with competitor names in some of them so share of voice has a denominator.
What is a good LLM visibility score?
There is no standard scale, and a single blended score is the wrong thing to chase. Track presence, share of voice, sentiment and prompt coverage separately, because each one has a different fix.
Can I trust a supplier’s LLM visibility report?
Only if they store full answer text, version their prompt set, record platform and date per run, and state their aggregation method. Without those, the trend is unfalsifiable.
Is a 400% improvement in three weeks believable?
Almost never. On a thirty-prompt set that is a handful of mentions, which is well within normal variance. Ask to see the stored answers from both runs — a genuine improvement survives that request.
Can any tool tell me why a model chose my competitor?
No. No provider exposes model reasoning, so every explanation is inference. Inference presented as measurement is the main integrity problem in this field.
What moves LLM visibility fastest?
Unblocking AI crawlers. It costs nothing, takes minutes, shows in server logs within days, and it is binary — blocked means absent regardless of anything else.
What moves it most in the long run?
Third-party corroboration: independent sources repeating what you say about yourself. Slowest to accumulate, largest durable effect, and impossible to fake convincingly.
Does publishing more content improve LLM visibility?
Barely. Retrieval selects passages rather than rewarding volume. Restructuring existing pages so each section answers its own heading returns far more than adding new ones.
How long until LLM visibility improves?
Crawler access shows within days, rendering in a week or two, extraction quality over weeks, presence on niche prompts in month two or three, and share of voice on competitive prompts over three to six months.
What is the difference between a citation and a mention?
A citation carries a link back to you; a mention names you without one. Both have value and they should be counted separately rather than merged.
What is prompt coverage?
Which buyer questions you appear on at all, and which you are entirely absent from. It is where the actionable gaps live and it is the measurement most often skipped.
What is share of voice in LLM answers?
How often you are named relative to the competitors who also appear across your prompt set. It needs competitor names in some prompts or it has no denominator.
Is LLM visibility the same as AI visibility?
In practice yes. ‘LLM’ is the technical framing and ‘AI’ the marketing one. The LLM-prefixed search terms are notably less contested while carrying comparable or higher cost per click.
Is LLM visibility the same as AEO?
No. LLM visibility is the metric; answer engine optimization is the practice that moves it. One is what you measure, the other is what you do.
Can I measure this for free?
Yes. Thirty prompts, two or three platforms, three to five runs each, answers pasted verbatim into a dated document, repeated monthly with identical wording. That is a real programme and it costs an hour a month.
Why must the prompt wording be fixed?
Because changing a single word changes which sources are retrieved. If this month’s prompts differ from last month’s, you do not have a trend, you have two unrelated snapshots.
What should I record on every run?
Prompt text, prompt set version, platform, region, date and time, run number, the full answer text, and the aggregation rule you decided in advance.
Does region affect LLM visibility?
Materially. Answers are location-sensitive, so region has to be held constant or month-to-month results are not comparable.

Want this done for your site?We build and maintain the search, content and paid programmes described on this page.

Get a free proposal

Get a free marketing proposal

Tell us what you are trying to grow and we will come back with a plan, not a pitch deck. Same-day reply on weekdays.

Privacy Preferences
When you visit our website, it may store information through your browser from specific services, usually in form of cookies. Here you can change your privacy preferences. Please note that blocking some types of cookies may impact your experience on our website and the services we offer.
Contact Us