Updated September 2026 · Written and maintained by the Progression Agency strategy team
Most sites fail before any content question arises — the crawler is blocked, or the page is empty without JavaScript. This is the technical half of AI visibility, written for the person who has to implement it.
On this page · 12 sections
- What is LLM SEO?
- What has to be true before an LLM can quote you?
- What does a technical LLM SEO audit actually check?
- Which AI crawlers should you allow?
- Which levers move which stage?
- What are the most common technical mistakes?
- What fixes actually work?
- What order should the work happen in?
- What should you ask a provider about the technical side?
- What do you deliver?
- How is technical LLM SEO measured?
- Does LLM SEO replace ordinary SEO?
The short answerLLM SEO is the work that makes a site retrievable and quotable by large language models. Four things have to be true in order: the crawler can reach the page, the content parses without JavaScript, the passage is selected as a candidate, and the wording is clear enough to reuse. Fail the first and nothing downstream matters — which is exactly where most of the sites we audit are failing, and it is the cheapest part to fix.
Updated September 2026. Crawler names and behaviour change; the technical checks here are re-verified quarterly.
What is LLM SEO?
LLM SEO is the technical and structural work that makes a website retrievable and quotable by large language models — so an assistant can reach your page, parse it, select it, and reuse your wording rather than a competitor’s.
It is the engineering half of what is also sold as answer engine optimization and generative engine optimization. The three terms describe substantially the same practice; LLM SEO is simply the name that gets used when the conversation is with a developer rather than a marketer. We explain that overlap properly on our terminology page.
The reason it deserves its own treatment is that the technical failures are both the most common and the least discussed. Agencies sell content programmes; the sites we audit are usually failing before any content question arises.
What has to be true before an LLM can quote you?
Four things, in order: the crawler can reach the page, the content parses without JavaScript, the passage is selected as a candidate, and the wording is clear enough to reuse.
Fail the first and nothing downstream matters. Most of the money spent on this discipline is spent on the fourth while the first is quietly broken.
Reach
The crawler has to be permitted and the server has to answer. Accidental blanket disallows, bot-management rules and aggressive rate limiting are the three commonest causes, and all three are usually invisible from the browser.
Parse
The content has to exist in the HTML. Many AI crawlers execute little or no JavaScript, so a page that assembles its text client-side can return effectively empty to the systems you care about while looking perfect to you.
Retrieve
Your passage competes against every other candidate for the question. Relevance to the question as asked matters more than keyword presence, because matching happens on meaning.
Reuse
The passage has to be quotable. Hedged, meandering writing gets paraphrased into something generic and uncredited; a crisp declarative sentence gets lifted intact.
What does a technical LLM SEO audit actually check?
Twelve things, split between access (can a bot get the content) and structure (can a model use it once it has).
Every one of these is verifiable. None of them requires trusting a proprietary score.
robots.txt
Which AI crawlers are permitted. We find accidental blocks constantly, usually inherited from a template or added during a scraping scare and never reviewed.
Server logs
The only unarguable evidence of which crawlers arrived and what they fetched. This is the first artefact we ask for and the one most clients have never looked at.
JavaScript rendering
What each key page contains with scripts disabled. The gap between that and what you see in a browser is the gap the crawler experiences.
Response time
Whether fetches complete before the crawler gives up. Slow pages are not merely ranked lower here; they are sometimes not retrieved at all.
Canonical consistency
Whether duplicate URLs split the signal across variants, which dilutes everything downstream.
Status codes for bots
Whether key pages return 200 to crawlers as well as to browsers. Bot management sometimes serves a challenge page instead, which reads as empty content.
Schema accuracy
Structured data that agrees with the visible page. Markup asserting something the copy contradicts reduces machine confidence in the rest of the page.
Heading structure
Whether each section answers one question. Headings phrased as questions, with the answer immediately beneath, are measurably easier to extract.
Answer position
Whether the answer precedes the explanation. This single change moves more than any other editorial intervention we make.
Entity consistency
Whether the same description of the organisation appears everywhere, without variation. Inconsistency is why assistants describe companies wrongly.
Internal linking
Whether related pages reinforce each other, which affects both retrieval and the model’s sense of what you are authoritative about.
Freshness signals
Whether update dates are genuine and visible. Fabricated freshness is detectable and counterproductive.
| # | Check | Evidence used | Typical failure |
|---|---|---|---|
| 1 | robots.txt and bot management | Live file plus WAF rules | AI agents blocked by a blanket rule |
| 2 | Server log crawler access | 90 days of raw logs | No AI user agent has ever arrived |
| 3 | JavaScript-off rendering | Fetch without JS execution | Page returns a shell and no text |
| 4 | Response time under load | Timed fetches from the bot’s perspective | Retrieval times out before the fetch completes |
| 5 | Canonical consistency | Rendered head versus sitemap | Conflicting canonicals split the signal |
| 6 | Status codes served to bots | Requests with each user agent | Bots served 403 while browsers get 200 |
| 7 | Schema accuracy | Markup compared to visible copy | Markup claims what the page does not say |
| 8 | Heading structure | Rendered heading tree | Headings used for styling rather than structure |
| 9 | Answer position | First sentence after each heading | The answer sits in paragraph four |
| 10 | Entity consistency | Descriptions across site and third parties | Three different self-descriptions |
| 11 | Internal linking | Crawl of the rendered site | Key pages reachable only through search |
| 12 | Freshness signals | Dates in markup versus real edits | Dates updated without content changing |
Which AI crawlers should you allow?
At minimum GPTBot, OAI-SearchBot, PerplexityBot, Google-Extended, ClaudeBot and Bingbot — unless you have a deliberate business reason to exclude one.
This is a commercial decision dressed as a technical one. Blocking protects your content from uncredited use; it also guarantees you are not cited. Most businesses selling something want to be found, and block by accident rather than by choice.
| Crawler | Operated by | What allowing it affects | Blocking cost |
|---|---|---|---|
| GPTBot | OpenAI | Training and retrieval for ChatGPT | High — largest assistant audience |
| OAI-SearchBot | OpenAI | ChatGPT’s search surface specifically | High |
| PerplexityBot | Perplexity | Live retrieval and citation | High — most citation-dense |
| Google-Extended | Google’s AI surfaces, separate from Googlebot | Moderate to high | |
| ClaudeBot | Anthropic | Anthropic model access | Moderate |
| Bingbot | Microsoft | Bing index, which feeds Copilot | High — also affects Bing |
Blocking Google-Extended does not remove you from Google Search. It is a separate control from Googlebot, and confusing the two is a common and expensive mistake in both directions.
Which levers move which stage?
Crawler rules move reach. Server rendering moves parsing. Answer structure moves reuse. Entity clarity moves retrieval. Nothing moves all four.
This is why single-tactic programmes stall. A schema-only engagement improves one column; an content-only engagement improves another; the site with a blocked crawler improves nothing at all until that is fixed.
What are the most common technical mistakes?
Blocking everything, rendering client-side only, schema that contradicts the page, burying the answer, treating llms.txt as a strategy, and publishing volume instead of clarity.
Five of these six are free to fix. The sixth is free to stop doing.
Blocking everything
Usually accidental. A rule added during a scraping scare, or inherited from a starter template, and never revisited.
Client-side rendering only
The most expensive silent failure. The page looks complete in a browser and returns nothing useful to a crawler that does not run scripts.
Schema that contradicts the page
Worse than no schema. It undermines machine confidence in everything else the page asserts.
Burying the answer
The commonest editorial failure. The answer exists, four hundred words down, and a competitor’s one-sentence version gets quoted instead.
Treating llms.txt as a strategy
Adoption is limited and it is not a ranking factor. Add it if you like; do not mistake it for the work.
Publishing volume
Thin pages dilute. One page that is genuinely the best answer outperforms fifty that restate the consensus.
| Mistake | Effect on retrieval | Cost to fix | Time to show up |
|---|---|---|---|
| Blanket crawler block | Total absence from every AI surface | Minutes | Days |
| Client-side-only rendering | Page reads as empty to most crawlers | Configuration change, occasionally a template rebuild | Weeks |
| Schema contradicting the page | Undermines confidence in the whole page | Hours | Weeks |
| Answer buried below the fold of the section | Passage not selected as a candidate | Editorial pass per page | Weeks |
| Treating llms.txt as the strategy | No measurable effect either way | Already done; the cost is the opportunity | Never |
| Publishing volume instead of clarity | Dilutes the pages that could have been cited | Ongoing, and usually negative | Months |
Want this done for your site?We build and maintain the search, content and paid programmes described on this page.
What fixes actually work?
Answer first, render server-side, allow the right bots, write one entity description, match the markup to the copy, and measure against a fixed prompt set.
Six fixes. Four of them can be done in a week by a developer who already has access.
What order should the work happen in?
Access, then rendering, then structure, then schema, then entity, then measurement. Each step depends on the one before it.
Restructuring content before confirming the crawler can reach it is the most common way to waste a quarter, and it happens because content work is easier to sell than log analysis.
- Pull server logs and confirm which AI crawlers arrived in the last ninety days.
- Check robots.txt and any bot-management rules for accidental blocks.
- Render your twenty most important pages with JavaScript disabled and read what remains.
- Fix anything that returns effectively empty.
- Restructure those pages so each section answers its heading immediately.
- Reconcile schema with the visible copy on the same pages.
- Write one canonical entity description and apply it everywhere.
- Build a fixed prompt set and take a baseline measurement.
- Re-measure monthly against the same set.
- Re-check crawler behaviour quarterly, because it changes.
What should you ask a provider about the technical side?
Ask whether they want your server logs. Everything else in an LLM SEO pitch is assertion; the logs are the one artefact that settles the question.
Four follow-ups separate the people doing this work from the people describing it.
Will you render our templates without JavaScript, and show us the output?
This is the check that finds the expensive failure. A provider who does not run it is guessing about the single most common reason a site is invisible, and a screenshot of the rendered text costs them ten minutes.
Which crawler user agents will you test, by name?
A vague answer about ‘AI bots’ means no one has looked at a log. The names are public, the list is short, and each one behaves differently enough that testing one proves nothing about another.
What will you measure before you start, and how often afterwards?
Without a baseline taken against a fixed prompt set, any later claim of improvement is unfalsifiable. The prompt set should be written down and handed to you, not kept in-house.
What will you not do?
A provider who claims to influence every AI surface is overselling. Some surfaces are index-led and respond to conventional work; some are live-fetching and respond to technical work; none of them accept payment for placement, and the honest answer names the limits.
Want the audit rather than the checklist?
We run this as a fixed-scope technical audit and hand you the findings with the fixes prioritised.
What do you deliver?
Six artefacts: crawler access report, render audit, schema reconciliation, answer-first rewrites, entity file, and a prompt set with a baseline.
All six are documents you keep, not a dashboard you rent.
How is technical LLM SEO measured?
Crawler hits in server logs, render completeness, schema validity, and then the same prompt-set measures used across all of this work.
The technical measures are binary and fast; the visibility measures are slow and probabilistic. Reporting that mixes them without saying which is which is reporting designed to look better than it is.
| Measure | Type | How fast it moves | How certain |
|---|---|---|---|
| AI crawler hits | Technical | Days | Certain — it is in the logs |
| Render completeness | Technical | Immediate | Certain |
| Schema validity | Technical | Immediate | Certain |
| Mention rate | Visibility | Weeks to months | Probabilistic |
| Citation rate | Visibility | Weeks to months | Probabilistic |
| Referral sessions | Visibility | Months | Under-reported by nature |
Does LLM SEO replace ordinary SEO?
No. Assistants that lean on a search index cannot retrieve a page that is not indexed, so conventional SEO remains the foundation this work sits on.
The honest framing is that the technical bar has risen. Everything that mattered for search still matters, and three things now matter that did not: whether AI crawlers specifically are allowed, whether content survives without JavaScript, and whether your answer is extractable.
What carries over
Crawlability, indexation, site speed, internal linking, canonical hygiene and authority. All of it still applies.
What is new
AI-specific crawler permissions, render-without-JS as a hard requirement, and answer-first structure as an editorial standard.
What changes in priority
Page speed matters more, because retrieval has tighter timeouts than indexing does.
What to stop doing
Optimising thin informational pages for traffic that assistants now absorb before the click.
Getting found in search
AI, AEO and what is changing
- GEO vs SEO
- How AI will affect SEO
- AI automation agency
- AI automation cost
- AI customer service
- AI in digital advertising
- AI for email marketing
- AI for restaurant social media
- Generative engine optimization agency
- AI visibility best practices
- Answer engine optimization agency
- AEO vs GEO vs LLM SEO
- How to rank in ChatGPT
- AI visibility tools compared
- AI visibility audit
- AI SEO agency
- AI search statistics, traced
- AI SEO in New York
- AEO vs SEO: what transfers
- AEO pricing: what it costs
- AEO agency: how to choose one
- What is AEO? Defined
- AEO services: what you receive
- AEO audit: the full checklist
- AI search optimization: two engines
- AEO tools: which class you need
- LLM visibility: what it measures
- Why AI recommends your competitor
- AEO search landscape: original dataset
- How AI search works: the pipeline
- How to get cited by AI: nine moves
- AEO myths: twelve claims assessed
- How long AEO takes
- Local AEO: why it differs
- Enterprise AEO: the real blockers
- AI visibility trackers vs rank tracking
- Why AI visibility dropped
- AEO content writing: six rules
- AEO for small business: where to start
- AI answers vs organic search
- AEO experts: the six skills
- AEO pros and cons
- AEO and paid search together
- URLs and AI citation
- AEO for dentists
- AEO for doctors
- AEO for plastic surgeons
- AEO for chiropractors
- AEO for therapists
- AEO for med spas
- AEO for fintech
- AEO for plumbers
- AEO for roofers
- AEO for HVAC
- AEO for landscapers
- AEO for hotels
- AEO for wineries
- AEO for ecommerce
- Free tool: AI crawler access checker
- Free tool: llms.txt generator
- Free tool: schema vs copy validator
- Free tool: extractable content checker
- Free tool: AI visibility prompt builder
Paid media and lead generation
Websites and design
Choosing and working with an agency
Social, content and brand
By industry and by situation
Frequently asked questions
What is LLM SEO?
Is LLM SEO different from AEO?
Is LLM SEO different from GEO?
Does LLM SEO replace traditional SEO?
Which AI crawlers should I allow?
Does blocking Google-Extended remove me from Google Search?
Can AI crawlers read JavaScript?
How do I check whether AI crawlers are reaching my site?
Does schema markup help LLM visibility?
Does llms.txt matter?
Do backlinks still matter for AI visibility?
How does RAG affect whether I get cited?
What is the single most common technical failure?
What is the cheapest fix with the biggest return?
How fast do technical fixes show up?
How long does an LLM SEO audit take?
What does an LLM SEO audit cost?
Do I need to rebuild my site?
Does page speed matter for LLM SEO?
What is answer-first structure?
How many pages should I restructure first?
Should I add FAQ schema?
Does content freshness matter?
Can I do this in-house?
What tools do I need?
How do I measure whether it worked?
Why is my competitor cited and I am not?
Does this work for ecommerce?
Does this work for small sites?
What should I ask a provider about the technical side?
Is there a downside to allowing AI crawlers?
How often should this be re-checked?
Want this done for your site?We build and maintain the search, content and paid programmes described on this page.
Why does Perplexity matter disproportionately for LLM SEO?
Because it is the shortest feedback loop available. Perplexity fetches at question time rather than relying mainly on prior knowledge, and it attaches citations to most of what it uses. If you unblock PerplexityBot and fix rendering on a Tuesday, Perplexity is where you are most likely to see evidence first.
That makes it a diagnostic instrument as much as a channel. Its audience is smaller than ChatGPT's, but a change visible in Perplexity and nowhere else usually means the work is correct and the other surfaces have not refetched yet — which is a different situation from the work being wrong.
What does Perplexity need from your site?
The same six things every AI surface needs, with two of them weighted more heavily: live fetchability and passage clarity.
PerplexityBot allowed
The crawler must not be blocked in robots.txt or by a bot-management rule. This is the single most common reason a site is absent from Perplexity specifically.
Fast response
Live retrieval has tighter timeouts than scheduled indexing. A slow page is sometimes not fetched at all rather than merely ranked lower.
Content in the HTML
Live fetching does not help if what arrives is a shell awaiting JavaScript.
Quotable passages
Perplexity cites densely, which rewards sections that answer their own heading and stand alone.
Attributed claims
Specific, sourced statements are safer to quote than confident assertions, and Perplexity quotes a lot.
Freshness
Genuine update dates help on a surface that fetches at question time. Fabricated ones are detectable and counterproductive.
How does Perplexity differ from ChatGPT and Google's AI surfaces?
Perplexity fetches live and cites densely. Google's AI surfaces lean on the Google index, so conventional SEO moves them most. ChatGPT sits between the two and has the largest audience of the three.
The practical consequence is sequencing. Technical work shows in Perplexity first, in ChatGPT's search surface next, and in Google's AI surfaces on the timescale of conventional indexing. Reading all three as one number hides that, which is why they should be measured separately even when the work is shared.
Should you optimise for Perplexity specifically?
No — optimise for retrievability and let Perplexity be the early reading. There is nothing Perplexity rewards that the others penalise, so a Perplexity-specific programme would be the same work with a narrower audience attached to it.
How the three surfaces respond to the same work
| Perplexity | ChatGPT search | Google AI surfaces | |
|---|---|---|---|
| Fetches live at question time | Yes, heavily | Yes | Partly — leans on the index |
| Cites with links | Densely | Yes | Sometimes |
| Responds fastest to crawler fixes | Yes | Second | Slowest |
| Responds most to conventional SEO | Least | Middle | Most |
| Relative audience size | Smallest | Largest | Large, but passive |
| Best used as | An early diagnostic | The commercial target | A conventional SEO outcome |
Frequently asked questions
Does Perplexity SEO differ from LLM SEO generally?
No. The same six workstreams apply. Perplexity weights live fetchability and passage clarity more heavily, which makes it the fastest place to see whether the work landed.
Why do I appear in Perplexity but not ChatGPT?
Usually timing rather than a fault. Perplexity fetches live; other surfaces refetch on their own schedule. It generally means the work is right and has not propagated yet.
What blocks Perplexity specifically?
PerplexityBot being disallowed in robots.txt or blocked by a CDN or WAF rule, and pages too slow to be fetched inside a live retrieval timeout.
Does Perplexity cite more sources than ChatGPT?
In our observation it names and links more sources per answer, which is why it is the most useful surface for checking whether a change worked.
Should I build a separate Perplexity strategy?
No. There is nothing Perplexity rewards that the other surfaces penalise, so it would be identical work aimed at a smaller audience.
How do I check whether Perplexity can reach my site?
Server logs. Filter for PerplexityBot by name across ninety days and count the fetches — and check the URLs, not only the total.
Related
Get a free marketing proposal
Tell us what you are trying to grow and we will come back with a plan, not a pitch deck. Same-day reply on weekdays.
