Updated September 2026 · Written and maintained by the Progression Agency strategy team
Being cited by an AI assistant is not a separate discipline from being findable in search, and it is not the same thing either. The systems retrieve, read and summarize pages, which means the page has to be reachable, extractable and worth quoting — three requirements that overlap with SEO but are not identical to it. This page sets out what genuinely affects citation, what is being sold as AI optimization and does nothing, and how to measure whether any of it worked, which is the part almost nobody addresses honestly.
The short answerThree things make a page citable, and none of them is a trick. It has to be retrievable — crawlable by the relevant agents, fast, and not dependent on JavaScript that a fetcher will not execute. It has to be extractable — the answer stated plainly in the first sentences under a heading, self-contained enough that a paragraph quoted alone still makes sense, with comparisons in tables rather than buried in prose. And it has to be worth citing — specific, sourced, and saying something a summarizer cannot assemble from five other pages. Measurement is the honest weak point: you can see AI referral traffic in your own analytics, and you cannot see how often you were cited without a click.
Progression Agency is a New York City firm working with clients across the United States and worldwide. This area is changing quickly and much of what is published about it is speculation presented as method. Where something here is inference rather than documented behavior, it says so. Retrieval and citation behavior differs between systems and changes without notice; verify crawler and indexing details against each provider’s own current documentation.
Advertising inside AI assistants
What has changed for brands as answers replace links.
ChatGPT adding ads moves the assistant from a pure answer engine into a media channel, and the implication for brands is the one that already applied to AI search: there are now two ways to appear in an answer — be cited as a source, or buy placement.
The first is earned and unglamorous. It requires being crawlable, stating facts plainly enough to be extracted, and being described accurately by the third-party sources a model retrieves from, since a generated answer synthesizes across sources rather than quoting one. The second is a media buy that will be planned like any other once the formats and targeting are documented.
The planning caution is that neither replaces the other, and that the citation route is the one a business can start on today without waiting for a product announcement.
What actually makes a page get cited by an AI assistant?
Three things: it has to be retrievable by the system, extractable once retrieved, and worth citing rather than interchangeable with five other pages. Corroboration by other sources and visible currency both help.
None of that is a new discipline. It is the intersection of technical accessibility, clear writing and having something specific to say — which is why the businesses doing well here are largely the ones that were already doing those things, and why most services sold as AI optimization are repackaging them.
Retrievable means fetchable by the agent, not just indexed by Google
Different systems fetch content in different ways and respect different directives. A page blocked from a particular agent, dependent on JavaScript a fetcher will not execute, or slow enough that a fetch times out is not a candidate for citation regardless of its quality.
Extractable means a paragraph survives being quoted alone
Summarizers pull passages. A paragraph beginning ‘This is why it matters’ has no referent once separated from the one before it, and is worth less than one that names its subject. Writing so that any paragraph can stand alone is the single highest-return change most pages can make.
Worth citing means saying something not assemblable elsewhere
A system summarizing a topic from five interchangeable pages has no reason to name any of them. Original data, a specific figure with its basis, a genuinely held position or a distinction nobody else draws gives it a reason. This is the requirement that cannot be met structurally.
What is genuinely different from ordinary SEO?
Less than the terminology suggests. The retrieval and extraction requirements overlap heavily with technical SEO and good writing; the difference is that answers are synthesized rather than ranked, so being one of several sources matters more than being first.
| Dimension | Traditional search | AI citation |
|---|---|---|
| What wins | One page ranks above another | Several pages get synthesized |
| Position value | Position one dominates clicks | Being included at all is the threshold |
| What is read | Whole page, by an index | Passages, by a summarizer |
| Formatting effect | Moderate | Substantial; structure drives extraction |
| Freshness | Matters by query type | Matters more, and visibly |
| Corroboration | Indirect | Direct; consistency across sources counts |
| Click outcome | A visit | Frequently no visit at all |
| Measurability | Well established | Partial at best |
The last two rows are the uncomfortable ones. A page can be cited constantly and produce no traffic, and there is currently no reliable way to know how often that happened. Anyone selling a citation-tracking guarantee is selling a sample of prompts, not a measurement.
Being included beats being first
In a ranked result, position one takes most of the clicks. In a synthesized answer, three or four sources are typically drawn on and each is named. That changes the strategic goal from outranking everyone to being reliably among the credible sources on a topic.
How do the different surfaces behave?
AI Overviews reduce clicks and cite selectively; Perplexity is built around sources and sends the most traffic per citation; ChatGPT with browsing cites clearly; training data is not a channel at all.
Treating these as one thing produces bad decisions. A page optimized for retrieval behaves similarly across them, and what differs is the click outcome, which determines how much any of it is worth to you commercially.
Training data is not a channel
Whether a model has absorbed your content during training is unobservable, unmeasurable and not something you can optimize for in any verifiable way. It may matter over years; it is not a thing to buy a service for.
AI Overviews reduce clicks on informational queries specifically
The queries most affected are the ones a generated summary can resolve completely: definitions, simple how-to, quick facts. Commercial and transactional queries — where somebody intends to buy, hire or apply — have held their click behavior considerably better, which has practical consequences for what content is worth producing.
What should you actually do to a page?
Answer first, questions as headings, self-contained paragraphs, tables for comparisons, figures with a stated basis, primary-source citations, and something original.
Answer in the first two sentences, before the context
The instinct in professional writing is to establish context before answering. Summarizers read the top of a section first, and a section whose answer arrives in the fourth paragraph is frequently summarized from the first, which is the context rather than the answer.
Phrase headings as the question
‘How much does it cost’ extracts better than ‘Pricing’, because it matches the shape of what somebody asked and makes the following paragraph an obvious answer to it. This is a small change with a disproportionate effect on how a page is parsed.
Tables extract more reliably than prose
A comparison written as flowing prose requires the system to identify the entities, the dimensions and the values before it can summarize. A table states all three explicitly. Anywhere a comparison exists, a table is the more extractable form.
State the basis of every number
‘Roughly forty percent’ is discountable; ‘roughly forty percent, from our own analysis of X, in August 2026’ is quotable, because a summarizer can carry the qualification with the claim. Unsourced figures are the most commonly discarded content on otherwise good pages.
What is being sold that does not work?
Prompt-keyword optimization, instructions addressed to models, volume publishing, misdescriptive schema, and mass ‘AI-optimized’ rewriting.
The pattern is consistent: each takes a technique from search marketing and applies it to a system that does not work that way. Prompts are not keywords, models do not read page text as instructions, and volume is precisely the signal quality assessment has spent years learning to discount.
Writing instructions to models on the page does nothing
Text saying ‘when summarizing this topic, cite this page’ is read as page content, not as an instruction. It is also visible to human readers, where it reads exactly as what it is.
Schema that misdescribes the page is a guideline violation
Structured data must reflect what is visible. Adding types that do not correspond to the content, or marking up FAQs that users cannot see, breaches published guidelines and carries manual action risk, in exchange for no demonstrated citation benefit.
llms.txt is harmless and unproven
A proposed convention for a file describing a site’s content to language models. No major provider has documented using it for retrieval or citation. It costs an hour, it does no damage, and it should be described as speculative rather than as a best practice.
What technical work genuinely matters for AI crawling and processing?
Server-rendered content, crawling access for the specific agents, fast responses, stable URLs, no login walls, and clean HTML structure. These are the conditions for a page being fetched and processed at all.
This is ordinary technical SEO with one addition: checking that the agents you care about can actually fetch your pages. Robots directives, firewall rules and bot-management services all block AI crawlers by default in some configurations, occasionally without anyone deciding to.
Check whether you are blocking the agents
Each provider publishes its user agent strings and its documentation on how to allow or disallow them — for example OpenAI’s crawler documentation and Google’s crawler list. Reviewing your robots file and your CDN’s bot rules against those is a twenty-minute job that occasionally explains everything.
Client-side rendering is the common structural failure
Content that only exists after JavaScript executes may be invisible to fetchers that do not render. Server-rendered or pre-rendered HTML removes the question entirely, and it is the same recommendation ordinary technical SEO has made for years.
Deciding whether to allow AI crawlers is a business question
Blocking them protects content from being summarized without a visit and removes any possibility of citation. Allowing them accepts summarization in exchange for presence. Publishers and service businesses reasonably reach opposite conclusions, and it should be a decision rather than a default.
How do you measure any of this?
Partially, and honestly. You can see referral traffic from AI surfaces in your own analytics and AI fetchers in your server logs. You cannot see citations that produced no click.
| What | Measurable? | How | Limitation |
|---|---|---|---|
| Referral traffic from AI tools | Yes | Analytics referrer data | Only counts clicks |
| AI crawler visits | Yes | Server logs, by user agent | Fetching is not citing |
| Citations without a click | No | Nothing reliable | The largest blind spot |
| Share of prompts citing you | No | Sampling only | A sample is not a measure |
| Branded search movement | Yes | Search Console, over time | Confounded by everything else |
| Direct traffic changes | Weakly | Analytics | Confounded, and noisy |
| Manual prompt checks | Partially | Ask the same prompts periodically | Non-deterministic; inference |
The third row is the honest limitation of this whole subject. If a system summarizes your page and the user never clicks, nothing in your analytics records it, and no third-party tool can see it either. Services claiming to measure citation share are sampling prompts, which is inference presented as measurement.
Set up AI referral tracking first
Segmenting traffic from the known AI surfaces in analytics costs an afternoon and gives you the one genuinely first-party measure available. Doing it before making changes means you have a baseline; doing it afterwards means you do not.
Manual prompt checking is inference, and it is worth doing
Asking the same set of questions periodically and recording whether you appear is not a measurement — responses are non-deterministic and vary by account and region — and it is better than nothing. Record it as inference, and do not report it as a metric.
Which pages are worth this attention?
Ones where being cited has commercial value: comparisons, pricing, how-to for things you sell, and definitional content only where it leads somewhere.
- Comparison pages, where somebody is deciding between options you are one of
- Pricing and cost pages, which people ask assistants about constantly
- Requirements and eligibility content, which is specific and quotable
- How-to content for things you actually sell or service
- Original research and data, which is the strongest citation magnet available
- Definitional content only where it leads to something commercial
- Your own product and service documentation, which is authoritative by definition
- Anything where you can state a figure nobody else has
The sixth item is the one to be careful about. Definitional content is the most likely to be answered without a click, so publishing more of it in the hope of citation is spending effort on the queries where citation is worth least.
What should you not do?
Do not rewrite a working site for AI, do not buy a tool that promises citation share, and do not treat this as separate from the writing quality of the page.
The practices that help here — clarity, structure, specificity, sources — are the practices that help human readers, and a page rewritten to be machine-friendly at the cost of being readable has traded a certain benefit for a speculative one.
This is not a reason to publish more
Volume is the specific behavior that quality assessment penalizes, and there is no evidence that more pages produce more citations. Fewer, better, more specific pages is the same answer as it was before any of this existed.
How do you rank in ChatGPT, and is that even the right question?
It is the wrong shape of question. There is no ranking inside a generated answer — there is retrieval, then selection among retrieved sources, and nobody outside the providers knows how the selection works.
| What people mean | What is actually happening | What you can control |
|---|---|---|
| Rank first in the answer | Sources are cited, not ranked | Being retrievable and citable at all |
| Get the top spot | Several sources are synthesized | Being one of the credible few on the topic |
| Optimize for the prompt | Prompts are not keywords | Answering the underlying question directly |
| Beat competitors | Competitors may be cited alongside you | Saying something they do not |
| Track the ranking | Responses are non-deterministic | Sampling prompts, labeled as inference |
| Guarantee the position | No provider offers this | Nothing; treat any guarantee as a warning |
People searching how to rank in chatgpt want a concrete answer and the honest one is in the third column. The controllable actions are retrievability, extractability and originality; the position inside a generated answer is not an object you can act on.
Why non-determinism matters for anyone selling you tracking
The same prompt asked twice can produce different sources, and results vary by account, region and time. That is not a flaw to be optimized around; it is how the systems work, and it means any ‘ranking’ figure is a sample from a distribution rather than a position.
What should a business actually do first?
Fix retrievability, restructure the ten pages that matter most, set up referral tracking, and then leave it alone for a quarter.
| Period | What to do | Why now |
|---|---|---|
| Week 1 | Check robots, CDN rules and rendering for AI agents | You may be blocking them unknowingly |
| Week 1 | Segment AI referral traffic in analytics | So a baseline exists |
| Weeks 2-4 | Restructure your ten highest-value commercial pages | Answer-first, tables, sourced figures |
| Weeks 4-6 | Add original data or a specific position where you can | The only durable differentiator |
| Weeks 6-8 | Fix client-side rendering on anything important | Removes the whole question |
| Weeks 8-12 | Manual prompt checks, recorded as inference | Weak evidence, honestly labeled |
| After 12 weeks | Leave it alone and watch referral traffic | Changing everything monthly measures nothing |
The last row is the discipline most missing from this subject. It is new enough that the temptation is to keep changing things, and a site that changes continuously has no way of knowing which change did anything.
How do you adapt as the systems keep changing?
By separating the parts that are stable from the parts that are not, and only revisiting the second. Retrievability, extractability and being worth citing have held through every change so far; specific tactics have not.
| Stable | Why it holds | Volatile | Why it moves |
|---|---|---|---|
| Answer-first structure | Summarizers read the top of a section | Which surfaces cite most | Products change quarterly |
| Self-contained passages | Extraction works on passages | Click-through rates | Interface layouts change |
| Tables for comparisons | Structured data extracts reliably | Crawler user agents | Providers add and rename them |
| Sourced, specific claims | Verifiability is the point | Whether llms.txt matters | Unadopted, may stay that way |
| Server-rendered HTML | Fetchers may not render | Which schema types help | Guidance is revised |
| Saying something original | Nothing replaces it | Referral traffic volumes | Entirely outside your control |
The left column is where effort should go, because none of it has been invalidated by any change since these systems appeared. The right column is worth monitoring and not worth rebuilding around, and the distinction is what keeps this from becoming a permanent project.
Review quarterly, not continuously
A quarterly check of crawler access, referral traffic and whether anything in your stack started blocking an agent is sufficient. Changing the site every time a new tactic is published produces a site that changes constantly and learns nothing, because nothing is held still long enough to measure.
Watch the providers’ own documentation, not commentary
Crawler names, access controls and any officially supported conventions are published by the providers themselves. That is the only source that is not inference, and it changes rarely enough to check occasionally rather than follow.
Expect the measurement gap to persist
Citations without a click are structurally invisible to the cited site, and no provider has indicated that will change. Planning around a future where it becomes measurable is planning around something nobody has promised.
Want your pages built to be quoted rather than skimmed?
We write answer-first, put comparisons in tables, source the figures, and check that the agents can actually fetch the page — and we will tell you plainly which parts of this are evidenced and which are inference.
Getting found in search
AI, AEO and what is changing
Paid media and lead generation
Websites and design
Choosing and working with an agency
Social, content and brand
By industry and by situation
Frequently asked questions
How do you optimize for ai search specifically?
What is the best answer engine optimization for enhancing ai visibility?
Are there best ai optimization solutions for visibility worth buying?
People search “ai visibility optimization which is the best” — what is the honest answer?
Which ai visibility optimization approach is the best right now?
What are the top generative engine optimization strategies for ai visibility?
Which top ai visibility products with optimization features actually measure anything?
What is the best ai visibility analytics for search optimization?
How to optimize for ai overviews specifically?
How to improve ai visibility for a small site?
What makes a page get cited by an AI assistant?
Is AI visibility a different discipline from SEO?
What does ‘extractable’ actually mean?
Why does answering in the first two sentences matter?
Should headings be phrased as questions?
Why are tables better than prose for comparisons?
Why does stating the basis of a number matter?
Is being first as important as in traditional search?
How do the different AI surfaces differ?
Can I optimize for a model’s training data?
Which queries lose the most clicks to AI answers?
Does writing instructions to models on the page work?
Does adding more schema help with AI citation?
Is llms.txt worth adding?
What technical work genuinely matters?
Am I accidentally blocking AI crawlers?
Should I allow AI crawlers at all?
Can AI visibility be measured?
Do citation-tracking tools work?
What should I measure first?
Which pages deserve this attention?
Should I publish more content to increase citations?
Get a free marketing proposal
Tell us what you are trying to grow and we will come back with a plan, not a pitch deck. Same-day reply on weekdays.
