Skip to main content Scroll Top

AI Visibility Trackers: Why Rank Tracking Cannot Do This Job

Updated September 2026 · Written and maintained by the Progression Agency strategy team

Two measurement approaches that barely overlap — and the one question that decides whether any number a tracker gives you can ever be audited.

On this page · 12 sections
  1. A rank tracker cannot measure AI visibility
  2. How prompt-set measurement works
  3. The four things to count, and why never to merge them
  4. When to buy a tracker, and when not to
  5. What to ask any tracker vendor
  6. What the tools cost
  7. Why rank tracking persists anyway
  8. Everything else we have written on search, AI and getting found
  9. Rank tracking versus prompt-set measurement
  10. The four metrics, and what a fall in each one means
  11. Verified tool prices, dated
  12. The free method, in full

The short answerA rank tracker cannot measure AI visibility — not poorly, at all. There is no ranked list inside a generated answer, and the failure is silent: a rank report looks healthy while you are absent from every answer your buyers receive. What replaces it is prompt-set measurement — a fixed, versioned set of buyer questions, re-run with identical wording, with every answer stored verbatim and four metrics counted separately. Below about thirty prompts you do not need a tool at all.

Tool prices loaded and read 16 September 2026. No product testing results are claimed here; this page compares measurement approaches rather than ranking products.

What a rank tracker can and cannot do hereWhat a rank tracker can and cannot do here
Three greens, five reds, and the five reds are the entire job. A rank tracker is not a weak AI visibility tool — it is measuring a different surface.

A rank tracker cannot measure AI visibility

Not poorly — at all. There is no ranked list inside a generated answer. Three or four sources are named, there is no second page, and no position to hold. A rank tracker is not a weak instrument for this job; it is measuring a different surface entirely.

This matters more than it sounds, because the failure is silent. A rank report can look perfectly healthy while the brand is absent from every answer its buyers actually receive. Nothing in the report says so.

We maintain several reviews of rank tracking tools and they remain useful for what they do. This page is about the thing they cannot do, and what replaces them.

Measures — Rank tracker. Position in a list of links.
Cannot measure — Rank tracker. Whether an answer named you.
Cannot measure — Rank tracker. Who was named instead.
Cannot measure — Rank tracker. How you were described.
Cannot handle — Rank tracker. Run-to-run variance.
Cannot store — Rank tracker. The answer text itself.
Two measurement approaches, compared on what they captureTwo measurement approaches, compared on what they capture
The two approaches barely overlap. The only shared column is auditability, and prompt-set measurement only earns it if verbatim answers are stored.

What a rank tracker does well

Position in a list of links, impressions and clicks from search, and fast detection of a ranking drop. All genuinely useful, none of it about answers.

What it cannot do

Tell you whether an answer named you, which competitors were named instead, or how you were described. Those three are the entire job.

Why the gap is invisible

Rankings and citations are different surfaces with different retrieval. A site can hold position one and be absent from every AI answer in its category, and no rank report will ever indicate it.

What replaces it

Prompt-set measurement: a fixed set of buyer questions, re-run with identical wording, with every answer stored verbatim.

How prompt-set measurement works

Fix a set of buyer questions in real phrasing and version it. Hold platform, region and account state constant. Run each prompt three to five times in one sitting, because variance is structural. Store every answer verbatim, dated and platform-labelled. Then count four things separately.

No tool is required for any of this. A document and an hour a month is a genuine measurement programme, and it is more auditable than several paid dashboards because the full answer text survives.

How prompt-set measurement actually worksHow prompt-set measurement actually works
No tool is required for this. Tools automate it past about thirty prompts or three platforms, which is where doing it by hand stops being realistic.
Presence — Prompt set. Named at all, across fixed questions.
Share of voice — Prompt set. You versus competitors named.
Sentiment — Prompt set. Whether the description is accurate.
Coverage — Prompt set. Which questions you appear on.
Stability — Prompt set. Consistency across repeated runs.
Auditability — Prompt set. Only if answers are stored verbatim.
Thirty questions — Free method. Buyer phrasing, versioned.
Two platforms — Free method. Held constant.
Three to five runs — Free method. Same sitting.
Paste verbatim — Free method. Dated document.
Count four things — Free method. Separately.
Monthly — Free method. Identical wording.

The four things to count, and why never to merge them

Presence, share of voice, sentiment and prompt coverage. Each has a different cause and a different fix, so a single blended score makes a decline impossible to diagnose.

If presence falls, the problem is usually access or retrievability. If share of voice falls while presence holds, it is structure or corroboration. If sentiment is wrong, it is entity data. One number tells you none of that.

Presence

Are you named at all, across the fixed set. The blunt first measure and the one most sites fail outright.

Share of voice

How often you are named relative to competitors who also appear. Needs competitor names in some prompts or it has no denominator.

Sentiment and accuracy

When you are named, is the description correct. Being named alongside a wrong claim is worse than not being named.

Prompt coverage

Which questions you appear on and which you are absent from entirely. This is where the actionable gaps are, and it is the measure most often skipped.

When to buy a tracker, and when not to

Buy past roughly thirty prompts, across three or more platforms, when somebody is committed to acting on the findings. Below that, a document and a monthly hour does the same job more rigorously.

The expensive mistake is not under-tooling. It is a subscription nobody acts on — bought to satisfy a reporting requirement, watched for two months, then ignored while the charge continues.

Manual, tool, or neitherManual, tool, or neither
Bottom-right is the expensive mistake: a subscription nobody acts on. It is more common than under-tooling.
Buy when — Past ~30 prompts. Manual stops being realistic.
Buy when — Three or more platforms. Coverage is the real product.
Buy when — Multiple markets. Prompt sets multiply.
Buy when — Someone will act. Otherwise it is a dashboard.
Do not buy when — Free checks not yet run. They find more.
Do not buy when — Rendering is broken. Measuring an unreadable page.

Run the free checks first

If robots.txt blocks a crawler or your pages need JavaScript to render, no tracker will tell you anything more useful than curl already would.

Then build a manual baseline

Thirty questions, two platforms, stored verbatim. An hour, and it is more than most organisations have.

Confirm someone will act

This is the question that decides whether the subscription is a programme or a dashboard.

Then automate

Coverage across platforms is the real product you are buying, not the interface.

What to ask any tracker vendor

Six questions, and one of them separates auditable measurement from unauditable counting: does it store verbatim answers, or only citation counts?

Because answers vary between runs, a count with no stored text cannot be verified by you, by us, or by the vendor. It is the question most buyers skip and the one that determines whether anything the tool reports can ever be checked.

What to ask before subscribing to any trackerWhat to ask before subscribing to any tracker
The verbatim-storage question is the one most buyers skip and the one that decides whether any number it produces can ever be audited.
Platforms named — Ask. Not 'all major AI engines'.
Verbatim storage — Ask. Counts cannot be audited.
Wording locked — Ask. Drift invalidates the trend.
Runs per prompt — Ask. And the aggregation rule.
Raw export — Ask. Or you are renting your data.
Published pricing — Ask. Several vendors moved to quote-only.
Explains model reasoning — Refuse. No provider exposes it.
Single blended score — Refuse. Hides which platform failed.
Guaranteed improvement — Refuse. Nobody controls attribution.
Rank data as AI proof — Refuse. Different surface entirely.
Counts with no text — Refuse. Unauditable by anyone.
Unversioned prompts — Refuse. No trend exists.

What the tools cost

Verified on 16 September 2026: Rankscale.ai $20 a month, Hall $49, Peec AI 89 euros, Surfer SEO $95, AthenaHQ $295. Profound publishes no paid tier at all — its pricing page offers a free trial and custom enterprise pricing.

That last point matters because Profound’s $499 figure is still quoted widely in roundups. It was accurate once and it is not what the page says today. We checked rather than repeating it.

Verified monthly list pricesVerified monthly list prices
Profound shows zero because its public pricing page lists a free trial and custom enterprise pricing only. Roundups still quoting $499/month are out of date.
Measurement, in numbersMeasurement, in numbers
The first and last zeros bound the category: rank trackers cannot start the job, and no tool can finish it by explaining causation.

Why rank tracking persists anyway

Not dishonesty, in most cases. Every agency already has a rank tracking stack, it produces one clean number that fits on a slide, clients already understand what position four means, and prompt measurement is genuinely messier — variance, repeated runs, four metrics and no single score to headline.

The decisive factor is that the failure is invisible. A rank report looks fine while you are absent from every answer, so nothing forces the change. Knowing that is usually enough to make it, which is the point of this page.

Why rank tracking persists as the defaultWhy rank tracking persists as the default
Worth stating plainly: most people reporting AI visibility with a rank tracker are not being deceptive. They are using the stack they already had.

Get a baseline that survives an audit

We build the prompt set in your buyers’ language, run it properly with repeated runs, store every answer verbatim and hand you the whole file. You keep it whether or not you work with us.

/ai-visibility-audit

Everything else we have written on search, AI and getting found

AI, AEO and what is changing

Websites and design

Choosing and working with an agency

Social, content and brand

By industry and by situation

Want this done for your site?We build and maintain the search, content and paid programmes described on this page.

Get a free proposal

Rank tracking versus prompt-set measurement

Two measurement methods compared on what each one can establish, what neither can, and why a rank tracker cannot be adapted to do this job.

Two instruments, two surfaces
QuestionRank trackerPrompt-set measurement
What does it observe?A list of linksA composed answer with named sources
Unit of resultPosition 1-100Named or not named
Is there a long tail?YesNo. Three or four sources, no second page
Handles variance?Not needed — results are stableEssential — repeated runs required
Stores the raw result?A position numberThe full answer text, if done properly
Tells you who beat you?Yes, by URLYes, by brand name
Tells you how you were described?NoYes
Can it be automated?FullyPast ~30 prompts, yes
CostExisting SEO stack$0 manual, or $20-$295/mo

The four metrics, and what a fall in each one means

Four things these tools report, and the diagnosis each decline points to.

Why a single blended score destroys diagnosis
MetricWhat it measuresIf it falls, the cause is usuallyThe fix
PresenceNamed at all across the setAccess, rendering or retrievabilityCrawler access, then rendering
Share of voiceYou versus competitors namedStructure or corroborationPassage rewriting, then digital PR
SentimentWhether the description is accurateEntity data disagreeing across sourcesOne canonical description everywhere
Prompt coverageWhich questions you appear onTopical gaps in what you have publishedTargeted content, not volume
Blended scoreNothing diagnosable—Do not use one

Verified tool prices, dated

Monthly list prices loaded and read from each vendor’s own pricing page rather than repeated from a listicle, each with the date it was checked.

Loaded and read 16 September 2026
ToolPublished priceBandNote
Rankscale.ai$20 / monthEntryCheapest verified
Hall$49 / monthEntry
Peec AIEUR 89 / monthProfessionalListed in euros
Surfer SEO$95 / monthProfessionalBroader SEO suite, not AEO-only
AthenaHQ$295 / monthEnterprise
ProfoundFree trial / customEnterpriseNo paid tier published; the widely-quoted $499 is stale

The free method, in full

The whole manual method written out, which produces the same underlying evidence for nothing.

What to do, and what it replaces
StepActionReplacesTime
1Write 30 buyer questions in real phrasingKeyword research1 hour, once
2Pick two or three platforms and hold them constantPlatform coverageMinutes
3Run each prompt 3-5 times in one sittingAutomated scheduling~1 hour
4Paste every answer verbatim into a dated docData storageIncluded above
5Count presence, share of voice, sentiment, coverageDashboard20 minutes
6Repeat monthly with identical wordingTrend reporting~1 hour/month

AI, AEO and what is changing

Frequently asked questions

Can a rank tracker measure AI visibility?
No. There is no ranked list inside a generated answer — three or four sources are named and there is no second page. A rank tracker is measuring a different surface entirely, and the failure is silent: a rank report can look healthy while you are absent from every answer.
What replaces rank tracking for AI?
Prompt-set measurement: a fixed, versioned set of buyer questions, re-run with identical wording across constant platforms, with every answer stored verbatim and four metrics counted separately.
What are the four metrics?
Presence (named at all), share of voice (versus competitors named), sentiment and accuracy (how you are described), and prompt coverage (which questions you appear on). Each has a different cause and fix, which is why merging them destroys diagnosis.
Why should I not use a single visibility score?
Because when it falls it cannot tell you which of the four moved. Presence falling points at access; share of voice falling points at structure or corroboration; sentiment points at entity data. One number points at nothing.
Do I need a tool to measure AI visibility?
Not below roughly thirty prompts on one or two platforms. A dated document and an hour a month is a genuine programme, and it is more auditable than a count-only dashboard because the full answer text survives.
When is a tracker worth buying?
Past about thirty prompts, across three or more platforms, and only when somebody is committed to acting on the findings. Coverage across platforms is the real product you are buying.
What is the most important question to ask a vendor?
Whether it stores verbatim answers or only citation counts. Because answers vary between runs, a count with no stored text cannot be verified by anyone — including the vendor.
How many times should each prompt be run?
Three to five, in one sitting, aggregated by a rule written down before you see the results. A source appearing once in five runs is not visible in any meaningful sense.
Why must prompt wording be locked?
Changing a single word changes which sources get retrieved. If this month’s prompts differ from last month’s you do not have a trend, you have two unrelated snapshots.
What do AI visibility trackers cost?
Verified on 16 September 2026: Rankscale.ai $20/month, Hall $49, Peec AI EUR 89, Surfer SEO $95, AthenaHQ $295. Profound publishes no paid tier — free trial and custom enterprise pricing only.
Is Profound still $499 a month?
Not according to its pricing page on 16 September 2026, which lists a free trial and custom enterprise pricing. The $499 figure persists in roundups and is out of date.
Can any tracker tell me why I was not cited?
No. No provider exposes model reasoning, so every explanation is inference. A tracker can establish that you are absent; only a diagnosis of your own site explains why.
Why do so many agencies still report AI visibility with rank data?
Mostly inertia rather than dishonesty. The stack already exists, it produces one clean number that fits a slide, clients already understand it, and the failure mode is invisible so nothing forces a change.
Should I keep my rank tracker?
Yes. It remains the right instrument for search positions, impressions and clicks. Just do not let it stand in for answer-engine measurement, and never merge the two into one report line.
How do I handle answers varying between runs?
Accept that it is structural rather than a bug, run each prompt several times, store every answer, and decide your aggregation rule before looking at results.
What should a monthly AI visibility report contain?
The four metrics separately, the platforms named individually, the prompt set version, and the stored answers available on request — reported alongside, never merged with, search rankings.
Can I export my data from these tools?
Ask specifically. If raw data cannot be exported you are renting your own measurements, and the history does not survive changing vendors.
How many prompts should I start with?
About thirty, covering definition, comparison, provider and problem intents, with competitor names in some so share of voice has a denominator.
Does the platform I test on matter?
Considerably. Coverage and reliability vary by platform, and answers differ materially between them. Test at least two and hold them constant month to month.
What is prompt coverage and why is it skipped?
Which buyer questions you appear on at all versus which you are entirely absent from. It is where the actionable gaps live, and it is skipped because it requires thinking about the set rather than reading a number.
Is manual measurement really as good as a tool?
At small scale it is arguably better, because the full answer text is preserved rather than reduced to a count. Tools win on scale, scheduling and multi-platform coverage, not on rigour.
What should I do first?
Run the free technical checks, then build a manual baseline of thirty prompts. Only buy a tracker once those are done and someone is committed to acting on what it shows.

Want this done for your site?We build and maintain the search, content and paid programmes described on this page.

Get a free proposal

Get a free marketing proposal

Tell us what you are trying to grow and we will come back with a plan, not a pitch deck. Same-day reply on weekdays.

Privacy Preferences
When you visit our website, it may store information through your browser from specific services, usually in form of cookies. Here you can change your privacy preferences. Please note that blocking some types of cookies may impact your experience on our website and the services we offer.
Contact Us