Skip to main content Scroll Top

Technical AEO: Crawler Access, Rendering and Verified Fixes for AI Search

Updated October 2026 · Written and maintained by the Progression Agency strategy team

Technical AEO is the engineering layer of answer engine optimization: the server, CDN, markup and URL work that lets the crawlers behind ChatGPT, Claude, Perplexity, Gemini, Microsoft Copilot and Google AI Overviews request a page, receive its full text and credit it to the right address. A technical AEO agency does that work as a service for companies whose sites sit behind a CDN or firewall, run on a JavaScript framework or span many templates and hosts, and it hands back each fix with proof from the server logs. Progression Agency is based in New York City and serves clients throughout the United States and worldwide.

On this page · 16 sections
  1. What is technical AEO?
  2. Which AI crawlers and user agents does the work cover?
  3. How does robots.txt apply to AI crawlers?
  4. Where do firewalls, CDNs and bot rules stop AI crawlers?
  5. Can AI crawlers read a page without JavaScript?
  6. What structured data belongs in the technical layer?
  7. How do canonical tags and URL hygiene fit in?
  8. Is llms.txt worth adding?
  9. How is crawler access verified in log files?
  10. How are the fixes verified?
  11. What does an engagement deliver?
  12. How long does the technical work take?
  13. How much does the technical layer cost?
  14. How do you choose a technical AEO agency?
  15. Is this the same as technical GEO or technical SEO for AI search?
  16. Related services

The short answerEvery major AI vendor documents its crawlers: OpenAI lists four user agents, Anthropic three and Perplexity two, while Google governs AI use through Googlebot and the Google-Extended token. The technical layer makes sure each agent you want is allowed in robots.txt, is not stopped by a firewall or bot rule, receives a 200 response with the text in the HTML, and finds one canonical URL per page with markup that matches what is visible. Each fix is written as a ticket with a pass-or-fail test, checked with a request that copies the crawler’s user agent, then confirmed in access logs against the vendor’s published IP ranges. An access fix is confirmed when the crawler next visits and is logged receiving the page. Whether an assistant then cites the page depends on content and corroboration, which the rest of an AEO program covers.

Agent names, stated purposes and controls are reported as each operator publishes them and were last checked on October 4, 2026. Operators add and rename agents, so the list is rechecked before every engagement. Passages marked editorial describe how we work. Ubersuggest’s United States data for October 2026 shows little or no search volume for this page’s phrases so far. Prices are the AEO planning bands this site publishes, and scoping in writing comes before any quote. No client is described and no result is claimed.

What is technical AEO?

It is the part of answer engine optimization that happens in configuration and code. It settles four things: whether the documented AI crawlers are permitted to request a page, whether anything between the internet and the server turns them away, whether the response contains the text, and whether the page declares one address and accurate markup for itself. Content, reputation and measurement are built on top of it and are described on our answer engine optimization agency page.

It is sold as a service, not handed over as a checklist, because the faults live in systems marketing teams rarely control: a CDN account, a web application firewall, a framework’s rendering mode, a deployment pipeline. The job is to find the fault, write the change so an engineer can ship it, and prove afterward that the crawler now receives the page.

How does it differ from technical SEO?

Technical SEO serves search engines that render JavaScript, publish detailed webmaster guidelines and report crawl problems in a console. The AEO layer serves a longer list of crawlers, most of which document only their names, purposes and IP addresses. The foundations are shared, and Google states that its AI features have no technical requirements beyond those of Search. The added work is access per vendor, tests that do not assume rendering, and verification in logs because there is no console to ask. Crawl budget, Core Web Vitals, migrations and index management stay with our technical SEO agency team.

How does it differ from the pages next to it?

Several pages on this site touch the technical side of AI visibility. Each has one job, and this page is the service that implements and proves the fixes.

Where each technical question is answered on this site
TopicThis pageCovered in full on
The item-by-item review a developer can followSummary onlyLLM SEO: the technical side
How retrieval and citation work inside an assistantNoHow AI search works
Whether URL length, depth or wording matterNo; only canonical and redirect fixes as service itemsURLs and AI citation
The audit as a fixed-price productNoAEO audit
WordPress sources of robots rules, plugins and cachingNoAEO for WordPress
Approval paths and release governance in large companiesNoEnterprise AEO
Crawler documentation by vendor, edge and firewall rules, rendering by framework, log verification and proof of fixesYesThis page

Who needs this as a service?

  • Sites behind Cloudflare, AWS WAF, Vercel or a host firewall whose bot settings nobody has reviewed since AI crawlers appeared.
  • Applications built with React, Vue or Angular where part of the content is assembled in the browser.
  • Sites with many templates, hosts or subdomains, where one robots.txt or one rule cannot be assumed to cover everything.
  • Teams that have been told ‘AI crawlers are allowed’ and want that shown in logs.
  • Companies about to change platform or CDN that want crawler access tested before and after the move.
  • Publishers and software companies that want to admit search crawlers while declining training crawlers, implemented exactly.

Which AI crawlers and user agents does the work cover?

Every agent an operator documents, grouped by what it is for. The operators publish three kinds: crawlers that build a search index, crawlers that collect training data, and fetchers that retrieve one page because a user asked. Each kind is controlled separately, and that separation is what makes a precise policy possible. The documentation to configure against is published by OpenAI, Anthropic, Perplexity and Google.

AI crawlers and fetchers as their operators describe them, October 4, 2026
OperatorAgent or tokenStated purposerobots.txt behavior statedIdentity check published
OpenAIOAI-SearchBotSurfaces sites in ChatGPT’s search featuresHonored; opting out removes the site from ChatGPT search answersIP range list
OpenAIGPTBotCrawls content that may be used to train foundation modelsHonoredIP range list
OpenAIChatGPT-UserVisits a page for a user action in ChatGPT or a custom GPTRules may not apply, because a user started the requestIP range list
OpenAIOAI-AdsBotChecks landing pages submitted as ChatGPT adsVisits only pages submitted as adsIP range list
AnthropicClaudeBotCollects web content that could contribute to model trainingHonored, including Crawl-delayPublished IP list
AnthropicClaude-SearchBotIndexes content to improve search resultsHonoredPublished IP list
AnthropicClaude-UserRetrieves a page when a Claude user asksCan be disabled in robots.txtPublished IP list
PerplexityPerplexityBotSurfaces and links sites in Perplexity search; not used for trainingControlled in robots.txtIP range list
PerplexityPerplexity-UserFetches a page to answer a user’s questionGenerally ignores robots.txtIP range list
GoogleGooglebotCrawls for Search, which AI Overviews and AI Mode are built onHonoredReverse DNS and IP range lists
GoogleGoogle-ExtendedToken only: governs use for Gemini training and groundingA robots.txt token, not a crawler; no effect on SearchNot applicable
GoogleGoogle-Agent and other user-triggered fetchersAct on a user’s requestGenerally ignore robots.txtIP range lists; Web Bot Auth in testing
MicrosoftBingbotBing’s crawler; Microsoft ties Copilot citations to Bing’s index and controlsHonored, with robots meta controlsIP range list
AppleApplebotSearch in Siri, Spotlight and SafariHonored; follows Googlebot rules if not named; ignores crawl-delayReverse DNS and IP range list
AppleApplebot-ExtendedToken only: opts content out of training Apple’s foundation modelsA robots.txt tokenNot applicable
AmazonAmazonbot, Amzn-SearchBot, Amzn-UserService improvement and possible training; search such as Alexa; user actionsFirst two honored; Amzn-User may not follow every directivePublished IP addresses
DuckDuckGoDuckAssistBotCrawls in real time for AI-assisted answers; not used for trainingHonored; an opt-out takes effect after 72 hoursPublished IP addresses
MistralMistralAI-User, MistralAI-IndexUser actions; indexing for Mistral search; neither used for trainingControlled with robots.txt tagsPublished IP addresses
Common CrawlCCBotBuilds an open repository of web crawl dataHonoredReverse DNS and IP range list

Three kinds of request: search, training and user-triggered

Search crawlers decide whether a site can be found and linked in an assistant’s answers: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot. Training crawlers and tokens govern model training: GPTBot, ClaudeBot, Google-Extended and Applebot-Extended. User-triggered fetchers retrieve one page because a person pasted a link or asked a question: ChatGPT-User, Claude-User and Perplexity-User. OpenAI states that each of its settings is independent of the others, so a site can admit search and decline training. The work starts by writing that decision down, because a rule aimed at training that also catches a search crawler takes the site out of answers.

Tokens that never appear in a log

Google-Extended and Applebot-Extended are not crawlers. Google describes Google-Extended as a standalone product token with no user agent string of its own: the crawling is done by Google’s existing agents, and the token controls whether content may be used to train Gemini models and to ground answers in Gemini Apps and Vertex AI. Google adds that the token does not affect inclusion or ranking in Search. Apple’s Applebot-Extended works the same way for Apple’s foundation models. A log review that waits for either name to show up is waiting for something that cannot happen.

How long does a change take to register?

Not instantly. OpenAI says its search systems can take about 24 hours to adjust after a robots.txt update, and Amazon gives the same figure for its agents. DuckDuckGo says an opt-out for DuckAssistBot takes effect after 72 hours. The protocol itself lets a crawler reuse a cached robots.txt for up to 24 hours. Google says excluding a site with its Search Console generative AI control generally takes a few days. Verification is scheduled around those intervals, so that a correct fix is not reported as a failure on day one.

Waiting times and limits that set the verification scheduleWaiting times and limits that set the verification schedule
Published operator and protocol figures, each named with its owner. They describe those systems, not any one website.

Want proof of what AI crawlers receive?Send the domain and what sits in front of it. We request your key templates as each documented crawler and return the status codes, response sizes and first tickets.

Request the access review

How does robots.txt apply to AI crawlers?

By the same standard as for any crawler, the Robots Exclusion Protocol published as RFC 9309. The parts that matter for AI access are how a crawler picks its group, what happens when the file cannot be fetched, and who else is writing to the file.

Which group does a crawler obey?

A crawler looks for the group whose user-agent line matches its product token, compared without regard to case, and obeys that group alone. The wildcard group applies only when no named group matches, and within a group the most specific matching path wins. Two consequences follow. A short group that names OAI-SearchBot and allows everything is not cancelled by a strict wildcard group further down. And a group that names a crawler in order to block one folder replaces the wildcard rules for that crawler, so anything the wildcard group disallowed is open to it again unless those lines are repeated.

What happens when robots.txt itself fails?

The standard separates two failures. If the file is unavailable, which for HTTP means a status in the 400 range, a crawler may access any resource. If it is unreachable because of a server or network error, a status in the 500 range, the crawler must assume everything is disallowed. Google documents its own handling: on a server error it stops crawling the site for the first 12 hours, then falls back on the last good copy of the file for up to 30 days. A robots.txt that returns a 500 or 503 for a few hours during each deployment is therefore an access fault, even though the contents of the file are correct. We test the file’s status from outside, repeatedly, including during a release.

Who else is writing the file?

The file a crawler receives is not always the file in the repository. Cloudflare’s managed robots.txt setting, when switched on, places Cloudflare’s own block ahead of the origin’s file, with disallow groups for named crawlers that include GPTBot, ClaudeBot, Google-Extended, CCBot and Amazonbot. Shopify generates the file itself and lets a theme template override it. Frameworks and CMS plugins produce theirs at build time or on request. The check is always made against the public URL of each host, and the result is compared with what the team believes it published. WordPress has its own set of sources, listed on AEO for WordPress.

Does Crawl-delay work?

For some agents. Crawl-delay is not part of RFC 9309. Anthropic says its crawlers support it; Apple says Applebot does not follow it. Where server load is the worry, Google’s guidance is not to answer with 401 or 403 as a way to slow crawling, and to treat 429 as the signal for too many requests. A page cache in front of the origin is the ordinary remedy for load, and it turns no one away.

Access policies and the robots.txt groups that express them
GoalAllowDisallowNote
Be cited by assistants and permit trainingEvery documented agentPrivate paths onlyThe simplest file; each agent is still tested
Be cited, decline trainingOAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot and the user fetchersGPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBotThe operators document these settings as independent
Stay in Google and Bing onlyGooglebot, BingbotThe named AI search crawlersPer OpenAI, the site then leaves ChatGPT search answers, though it may still appear as a navigational link
Keep one section out of answersEverything elseThat path, in every named groupNamed groups do not inherit wildcard rules, so the path is repeated
Reduce load without blockingAll agentsNothingUse caching, and Crawl-delay where it is supported; avoid 403

Where do firewalls, CDNs and bot rules stop AI crawlers?

In front of the server, where robots.txt has no say. A robots.txt rule is a request that well-behaved crawlers honor; a firewall rule is enforcement. The operators expect this layer to matter: OpenAI recommends allowing its published IP ranges as well as its robots.txt token, Perplexity publishes firewall instructions for Cloudflare and AWS, and Anthropic says its crawlers will not try to get past a CAPTCHA.

CDN — Bot policy at the edge. Search, agent and training traffic handled separately.
Firewall — Managed and custom rules. User agent, category, country and rate.
robots.txt — One file per host. Named groups, availability and who else writes to it.
Origin — Server and cache. Status codes, response size and freshness.
Application — Rendering mode. Whether the text is in the first HTML response.
Markup — Head and structured data. One canonical and a graph that matches the page.

Cloudflare

Cloudflare sorts AI traffic into Search, Agent and Training in its AI bot policies and lets each be allowed, blocked everywhere or blocked only on pages that carry ads. Its documentation gives September 15, 2026 as the date from which new domains default to blocking Training and Agent bots on pages with ads while leaving Search allowed, and notes that a crawler used for both search and training is caught by any option that blocks training. Bot Fight Mode issues computational challenges to traffic it identifies as bots and cannot be adjusted with custom firewall rules. AI Crawl Control reports which AI services request a site and sets allow or block rules per crawler. We read all three settings, zone by zone, before touching robots.txt.

AWS WAF

The AWS managed Bot Control rule group contains a rule named CategoryAI that inspects for artificial intelligence bots. Its documented default action is Block, and AWS notes that the action applies to verified and unverified bots alike. A site that enabled Bot Control against scrapers is therefore refusing the documented AI search crawlers unless that rule’s action was overridden. AWS also documents that Bot Control verifies bots by the origin IP address of the request, which matters when a proxy or load balancer sits in front and the true client address travels in a forwarded header.

Vercel and other platform firewalls

Vercel offers an AI bots managed ruleset that identifies known AI crawlers and can be set to log or deny, and a separate bot protection ruleset that serves a JavaScript challenge to non-browser clients while exempting verified bots. Other hosts and security products have equivalents under different names. The method does not change with the product: find every rule that matches on user agent, bot category, country or request rate, then test what each documented crawler receives.

Challenge pages, rate limits and country blocks

A challenge page is HTML that is not the content, often delivered with a 200 or 403 status. To a crawler that does not solve challenges, it is the page. Rate limits answer fast crawlers with 429 or 503. Country blocks stop crawlers whose IP ranges sit outside the permitted countries. Each of these shows in logs as a status or response size that differs from what a browser receives for the same URL, which is why verification relies on logs and not on screenshots.

Layers that can refuse an AI crawler, and what each is documented to do
LayerSetting or ruleDocumented behaviorWhat the service does
CDN bot policyCloudflare AI bot policiesSearch, Agent and Training are each allowed or blocked; new-domain defaults block Training and Agent on pages with adsSets each zone to match the written policy
CDN bot challengeCloudflare Bot Fight ModeChallenges identified bots; not adjustable with custom rulesAgrees on or off with the security owner; tests each crawler
Managed firewall rulesAWS Bot Control, CategoryAIDefault action Block, for verified and unverified botsOverrides the action or narrows the rule
Platform firewallVercel AI bots rulesetLog or deny for known AI botsSets it to log unless the policy is to block
IP allowlistsOperator IP range listsOpenAI and Perplexity recommend admitting their published rangesAutomates updates from the published lists
robots.txt at the edgeCloudflare managed robots.txtAdds disallow groups for named crawlers ahead of the origin fileTurns it off or reconciles it with the origin file

For a comparison of those two platforms themselves, see Cloudflare vs AWS.

Can AI crawlers read a page without JavaScript?

Assume they cannot, and build so that it does not matter. Google renders JavaScript. Most other operators say nothing on the subject, and Microsoft warns that AI systems may not render hidden content. The engineering target is that the text an assistant would quote is present in the HTML the server sends.

What do the operators say about rendering?

Google’s JavaScript documentation describes a rendering queue in which a headless Chromium executes scripts after the first fetch, and adds that server-side rendering or pre-rendering is still a good idea because not all bots can run JavaScript. Its note on dynamic rendering calls that technique a workaround, not a recommendation, and observes that other search engines may ignore JavaScript. The crawler documentation from OpenAI, Anthropic and Perplexity describes what each agent is for and how to control it; none of it states that the agent executes scripts. Apple says Applebot may render pages and advises that sites degrade gracefully when resources cannot be loaded. Microsoft’s guidance for AI search answers advises against hiding important answers in tabs or expandable menus, because AI systems may not render them, and against keeping key information only in images or PDFs.

What do the frameworks do by default?

It depends on the framework and, within a framework, on choices made per route. The table gives each one’s documented default and the place to look for content that has left the server response.

Rendering defaults as each framework documents them
FrameworkDocumented defaultWhat to check
Next.js (App Router)Layouts and pages are Server ComponentsClient Components that fetch the main content after load
React without a frameworkServer APIs can render components to HTML; a framework normally calls themWhether any server rendering is configured at all
AngularApplications are client-side rendered; server-side and hybrid rendering are opt-inWhether SSR or prerendering is switched on for public routes
VueComponents render in the browser; server rendering and static generation are availablePages that open on a loading state and then fetch their content
SvelteKitPages render on the server first, then hydrate; ssr, csr and prerender are page optionsRoutes where server rendering has been turned off
AstroMostly static HTML, with interactive islandsContent placed inside client-only islands

Our Next.js development and React development teams implement the change where a client has no engineers to spare. Headless WordPress and headless CMS vs traditional CMS cover decoupled front ends, and hosting for React and Next.js covers where server rendering runs.

The raw-HTML test

The test is small, repeatable and independent of any tool’s score. It is run per template, not per page.

  1. List the templates that carry commercial content: home, service or product, pricing, location, article, FAQ and comparison.
  2. For each template, choose one URL and three sentences an assistant would need in order to answer a buyer.
  3. Request the URL with a command-line client that runs no scripts: once with a browser user agent, once with each documented crawler’s.
  4. Search each response body for the three sentences.
  5. Record the status code, the response size and whether each sentence was found.
  6. Where a sentence is missing, identify the component that loads it and the data source behind it.
  7. Write the fix as server rendering, prerendering or moving the text into the first response, with the same three sentences as its test.

What usually goes missing?

  • Pricing tables filled from an API call after the page loads.
  • Product specifications inside tabs that fetch their panel on click.
  • Reviews and ratings injected by a third-party script.
  • FAQ answers revealed by a script instead of sitting in the markup.
  • Store, dealer and location finders that exist only as a map.
  • Content held back by a cookie or consent layer until someone clicks.
  • Text inside iframes, images and PDFs.
  • Views that exist only as URL fragments, with no address of their own.

Running React, Vue or Angular?Tell us the framework and hosting. We run the raw-HTML test on each template and show which content never reaches the server response.

Test my templates

What structured data belongs in the technical layer?

Enough to state plainly who publishes the page and what it describes, and nothing the page does not show. Structured data is confirmation for a machine. It is not a route into answers.

What do Google and Microsoft each say?

Google’s guide to generative AI features in Search says structured data is not required for them, that no special schema.org markup exists for AI, and that markup remains worth using for rich results provided it matches the visible text. Its structured data policies rule out marking up content readers cannot see and markup that misrepresents the page. Microsoft’s guidance says schema helps search engines and AI systems understand content, and recommends the JSON-LD format. The two positions fit together: accurate markup helps identification, and neither company treats it as a substitute for the text.

The graph we ship

  • Organization, with legal name, logo, address, contact point and links to verified profiles.
  • WebSite and WebPage, tying each URL to the site and to its breadcrumb.
  • Service, or Product with Offer, where the page describes something sold and states a price.
  • Article with author and dates on editorial pages, where dates change only when the content does.
  • FAQPage where the page shows real questions with their full answers.
  • BreadcrumbList on every page below the home page.

All of it is JSON-LD in schema.org vocabulary, one connected graph per page, generated from the same fields that produce the visible copy so that the two cannot drift apart.

Markup that contradicts the page

The fault we look for is disagreement: a price in the markup that differs from the price in the table, opening hours changed on the page but not in the plugin, a rating with no visible reviews, two plugins describing one page as two organizations. Each template is validated and then compared field by field with the rendered text. Our schema and copy validator automates the comparison, and the schema markup generator writes clean markup for hand-built sections.

How do canonical tags and URL hygiene fit in?

They decide which address gets the credit. Google’s generative AI performance report assigns most of its data to a page’s canonical URL, and Bing reports citations per URL, so a page reachable at four addresses splits its own evidence four ways.

One address per page

Google treats redirects and the rel=canonical link as strong canonical signals and sitemap inclusion as a weak one; the link relation itself is defined in RFC 6596. The service items are concrete: one protocol and host, one trailing-slash rule, parameters that create no indexable copies, redirects that reach the final address in one hop, and a sitemap that lists only canonical URLs with honest modification dates. Where the platform supports IndexNow, which Microsoft recommends for keeping AI answers current, changed URLs are submitted as they change.

Canonical in the HTML, not added by a script

Google’s JavaScript documentation says the best place to set a canonical is the HTML, and warns against using a script to change it to a value different from the one the server sent. A crawler that runs no scripts sees only the server’s version, so a canonical injected or rewritten in the browser is, for that crawler, missing or wrong. The same document advises routing single-page applications with the History API instead of URL fragments, so that each view has a real address.

  • Redirect chains shortened to a single hop.
  • http and https, www and bare-host variants resolved to one.
  • A canonical tag present in the server response on every template.
  • Parameterized duplicates canonicalized or kept out of the crawlable set.
  • Error pages that answer 200 changed to return the correct status.
  • Sitemap entries limited to canonical, indexable URLs that return 200.

Whether URL length, depth or wording matter is a separate question, answered on URLs and AI citation.

Is llms.txt worth adding?

It is an optional extra, added when it is cheap to generate and easy to keep accurate. It is not an access control, and it is not a ranking signal.

What does the proposal specify?

The llms.txt proposal describes a Markdown file, at the site root or at any path, that gives a short description of the site or section and links to the pages an agent would need, ideally to Markdown versions of them. Its current version also suggests link relations, in the HTML or in HTTP headers, that point from a page to its Markdown version and to the llms.txt file covering it. The format is deliberately plain: a title, a short summary and lists of links.

Who reads it?

Google states that Search does not use llms.txt, and that keeping one neither helps nor harms visibility there. The proposal’s own account is that the files are used most heavily for software documentation, where coding agents follow them to find references. The developer documentation sites of OpenAI and Perplexity point readers to an llms.txt index and to Markdown versions of their pages, which shows the convention at work for documentation. None of the crawler documentation we read says a search or training crawler gives the file special treatment.

When do we add one?

When the site has documentation, policies or product data that agents are likely to be sent to; when the CMS can generate the file from live content, so it never goes stale; and when every link in it resolves to a page that is indexable and current. We do not add one in place of fixing access or rendering. The free llms.txt generator produces a first draft.

How is crawler access verified in log files?

By finding the crawler’s own requests in the access logs and checking three things: that the request really came from the operator, what status and size the server returned, and which URLs were asked for. Logs are the only record of what a crawler was served. Every other test is a simulation.

Which logs, and which fields?

The log has to come from the outermost layer that answers requests. If a CDN sits in front, the origin’s log shows only what the CDN let through, and a crawler blocked at the edge never appears in it. We ask for 30 to 90 days from each layer: the CDN’s request logs at the edge, and the web server’s access log at the origin, whether that is Apache, NGINX or IIS. The fields needed are the timestamp, the requested host and path, the status code, the bytes sent, the user agent and the client IP address as seen at the edge.

How are real crawlers told from imitators?

A user agent string is a claim, not proof. Common Crawl itself warns that other crawlers falsely identify as CCBot. The operators publish stronger identifiers, and the review uses whichever one each offers.

How each operator lets you confirm that a request is genuine
OperatorMethod publishedHow the review uses it
OpenAIA list of IP ranges for each agentMatch the logged address against the list for the agent named
AnthropicA published list of source IP addressesMatch the logged address against the list
PerplexityA list of IP ranges for each agent, with firewall instructionsMatch the address; mirror the allow rule it describes
GoogleReverse DNS to googlebot.com or google.com hosts, plus IP range lists per crawler classReverse lookup, then forward lookup, or match the lists
AppleReverse DNS in the applebot.apple.com domain, plus a list of IP rangesReverse lookup or match the list
MicrosoftA list of Bingbot IP rangesMatch the logged address against the list
Amazon and DuckDuckGoPublished IP addresses for each agentMatch the logged address against the list
Common CrawlReverse DNS to crawl.commoncrawl.org, plus a list of IP rangesReverse lookup or match the list

A newer method is cryptographic. Web Bot Auth signs HTTP requests so that the receiver can verify the sender against a public key directory. Cloudflare uses it as one way to verify bots, AWS WAF labels requests by signature status, and Google says it is experimenting with the protocol for its Google-Agent fetcher. Where a client’s edge supports it, signature status becomes one more column in the review.

What should the log show after a fix?

  • Requests from the operator’s verified IP ranges carrying the expected user agent.
  • A 200 status on the URLs that matter, and 301 only where a redirect is intended.
  • Response sizes in line with what a browser receives for the same URL.
  • A robots.txt fetch from each crawler that returns 200.
  • No run of 403, 429 or 503 responses to a documented agent.
  • Requests that reach deep templates, not only the home page.

What if there are no logs?

Some platforms do not expose raw access logs. The evidence then comes from whatever the platform does provide: a CDN’s AI crawler report, a firewall event log, or a logging rule added for the purpose that records requests from the documented user agents. Where nothing can be captured, the report says so and rests on synthetic tests alone, labeled as such.

  1. Collect logs from every layer and every host for the same period.
  2. Filter for the documented user agents.
  3. Verify each matching request against the operator’s IP list or by reverse DNS, and set imitators aside.
  4. Tabulate by agent, host, status code and template.
  5. Compare response sizes with a browser’s for the same URLs.
  6. List the URLs requested and the URLs never requested.
  7. Read the robots.txt fetches separately.
  8. Write each anomaly as a finding, with the log lines attached.
The path of one crawler requestThe path of one crawler request
Editorial diagram of the layers reviewed in an engagement, in the order a request meets them.

Need the ranges before a call?AEO audits, retainers and implementation projects are published as planning ranges. Ask for the band that fits your hosts and templates.

See AEO pricing

How are the fixes verified?

In four levels of increasing strength, and a fix is not closed until it reaches the third. A recommendation is not a result, and a deployment is not a result either until the crawler has been served the corrected page.

Tickets with acceptance tests

Each finding becomes a ticket an engineer can act on without a meeting: the fault, the evidence, the change, the systems touched, the risk, and a test that passes or fails. ‘Allow AI crawlers’ is not a ticket. A usable one reads like this: OAI-SearchBot receives 403 on the pricing template from the edge; override the Bot Control CategoryAI action for requests from OpenAI’s published ranges; test: a request with the OAI-SearchBot user agent returns 200 and contains the three reference sentences.

Fixes, their acceptance tests and the evidence kept
FixAcceptance testEvidence filedRechecked
robots.txt group added or correctedThe public file shows the group on every host; a parser test returns allowed for the target URLsBefore and after copies of the file, with timestampsAfter each deploy
Firewall or bot rule changedA request with the crawler’s user agent returns 200 and the full bodyThe rule diff; response headers and sizeWeekly, and after any rule change
IP allowlist addedRequests from the operator’s ranges are not challengedEdge log lines from verified addressesWhen the operator’s list changes
Template server-renderedThe three reference sentences are present in the raw HTMLSaved response bodiesAfter each release
Canonical correctedThe server response carries one self-referencing canonicalA capture of the head for each templateAfter each release
Structured data reconciledThe validator passes and values equal the visible textValidator output and a comparison tableMonthly
Redirect chain shortenedOne hop to a 200Trace outputAfter URL changes
robots.txt availability200 on repeated fetches, including during a deployAn uptime record for the fileContinuously

Four levels of evidence

  1. Configuration: the rule, file or template has changed, shown by a diff.
  2. Synthetic request: a request copying the crawler’s user agent now receives the right status and body. This proves the path is open for that string, not that the operator’s network is admitted.
  3. Verified crawler request: the operator’s own crawler, confirmed by IP range or reverse DNS, is logged receiving a 200 with the full response.
  4. Downstream signal: the page shows as a cited URL in Bing Webmaster Tools, gains impressions in Search Console’s generative AI report, or is linked in a prompt-set answer.

Level three closes the ticket. Level four is reported when it happens and is never promised, because citation depends on content and competition as well as access.

Level 1 — Configuration changed. Shown by a diff of the rule, file or template.
Level 2 — Synthetic request passes. Correct status and body for the crawler's user agent.
Level 3 — Verified crawler served. Operator's own IP logged with a 200 and full size.
Level 4 — Downstream signal. Cited URL or impressions in a first-party report.
Closes a ticket — Level 3. The crawler itself was served the corrected page.
Never promised — Level 4. Citation depends on content and competition too.
What counts as proof that a crawler can read a page (editorial)What counts as proof that a crawler can read a page (editorial)
Editorial scorecard. Green closes a ticket, amber is a useful first pass, red proves nothing about what a crawler was served.

Regression checks after every release

Access breaks quietly: a new firewall rule, a framework upgrade that changes a rendering default, a CDN migration, a plugin that rewrites robots.txt. The service therefore leaves monitoring behind it: scheduled requests per template and per documented user agent, an alert when a status code or response size changes, a daily copy of each host’s robots.txt with an alert on any difference, and a recurring log review. Where the client has a deployment pipeline, the raw-HTML test for the reference sentences is added to it, so that a release which removes the text fails before it ships.

What does an engagement deliver?

Shipped changes, each with its proof, and the monitoring that keeps them in place.

Policy — Written access policy. Which agents are allowed for search, training and user fetches.
Log report — Verified crawler requests. By agent, host, status and template.
Layer map — Every rule that touches bots. CDN, firewall, platform and application.
HTML report — Raw response per template. Reference sentences found or missing.
Tickets — Changes with tests. Written so an engineer can ship them.
Monitoring — Checks that stay behind. Scheduled requests and a robots.txt watch.

Deliverables

  • A written access policy: which agents are admitted for search, for training and for user-triggered fetches.
  • A crawler access report per host, built from logs, with verified requests separated from imitators.
  • A layer map of every CDN, firewall, platform and application rule that touches bot traffic.
  • A raw-HTML report per template, with the reference sentences found or missing.
  • Tickets with acceptance tests, in priority order.
  • Implementation where access allows, or pairing with your engineers where it does not.
  • A reconciled structured data graph per template.
  • Canonical, redirect and sitemap corrections.
  • An evidence file for each closed ticket.
  • Monitoring and a recurring log review.

What is not included

  • Writing or restructuring page copy, which is AEO content writing.
  • Prompt-set measurement and reporting, which belong to the AI visibility audit.
  • Getting around an operator’s controls, or disguising requests.
  • Serving crawlers different content from the content users receive.
  • Guarantees of citation.
  • Security decisions: we recommend, and your security owner approves.

How long does the technical work take?

Plan on two to six weeks for a single site, from access to the logs to closed tickets, when changes can be deployed as they are approved. That is a planning range. The work itself is small; elapsed time is governed by log access, change approval and release schedules.

Planning timeline for one site
PhaseWhenWorkOutput
Access and inventoryWeek 1Logs, CDN and firewall access; host and template inventory; written policyLayer map and access policy
ReviewWeeks 1 to 2Log analysis, requests per agent, raw-HTML testsFindings with evidence
TicketsWeek 2Changes written with acceptance tests and priorityTicket list
ImplementationWeeks 2 to 5Configuration changes first; rendering and template changes afterDeployed fixes
VerificationFrom each deploy, plus the operator’s stated delaySynthetic tests at once; verified crawler requests as they arriveEvidence files
MonitoringOngoingScheduled tests, robots.txt watch, recurring log reviewAlerts and a short monthly note
Planning timeline for the technical workPlanning timeline for the technical work
Planning intervals, not promises. Elapsed time is set by log access, change approval and release schedules more than by the work.

Rendering changes on a client-rendered application can take longer than everything else combined, and are scoped separately after the raw-HTML test. Approval time in large organizations is a subject of its own, covered on enterprise AEO. Citation timing after access is fixed is a different question; how long AEO takes deals with it.

How much does the technical layer cost?

It is priced from the published AEO planning ranges. A review with tickets falls in the audit band; implementation is scoped after the review, because rendering work varies so widely. A quote follows a written scope.

Published AEO planning bands and the technical work each one holds (US dollars)
Planning bandEngagement typeTechnical work it holds
$1,000 to $4,000, one timeAEO auditThe review: logs, access per agent, raw-HTML tests and a ticket list (the band is published for 10 to 20 hours)
$25,000 to $100,000, per projectImplementationA full rendering and restructuring program on a client-rendered site
$1,500 to $5,000 per monthRetainer, small businessMonitoring and fixes as part of a wider AEO program
$5,000 to $10,000 per monthRetainer, mid-marketThe same, across more templates and platforms
$10,000 to $20,000 or more per monthRetainer, enterpriseThe same, across many hosts and teams

What moves a technical quote?

  • Rendering: a server-rendered site needs configuration; a client-rendered one may need engineering.
  • The number of templates, hosts and subdomains.
  • The number of layers in front of the origin, and who controls each.
  • Whether raw logs exist, and how far back they go.
  • Who implements: your engineers with our tickets, or our developers.
  • Whether monitoring continues inside a retainer.

The full set of ranges and what drives them is on AEO pricing and in the marketing agency pricing guide.

How do you choose a technical AEO agency?

Choose on evidence habits. The work is verifiable end to end, so a provider’s method shows in what they ask for on the first day and what they hand over on the last.

What to require of a provider, with the check for each
What to requireThe check
Works from logsThey ask for edge and origin logs before proposing fixes
Names crawlers preciselyThey separate search, training and user-triggered agents, operator by operator
Knows the operators’ documentationAsk what each operator says its agent is for and how it is controlled
Tests without renderingThey show raw HTML responses, not browser screenshots
Verifies identityThey check IP ranges or reverse DNS, not user agent strings alone
Writes testable ticketsEach ticket carries a pass-or-fail test
Proves fixesThey return log lines from verified crawlers after deployment
Leaves monitoring behindScheduled tests and a robots.txt watch remain when they finish
Respects security ownershipFirewall changes are proposed to your security team, not made around it

What should you ask before signing?

  1. Which crawlers will you test, and what does each one’s operator say it is for?
  2. Which logs do you need, from which layers, and covering how long?
  3. How will you tell a real crawler from a request that copies its user agent?
  4. What do you do when our platform exposes no logs?
  5. How will you test rendering without a browser?
  6. What does one of your tickets look like?
  7. What evidence closes a ticket?
  8. Who makes the change in the CDN, the firewall and the application, and who approves it?
  9. What monitoring stays in place, and who receives the alerts?
  10. Which parts of this are one-time, and which need a retainer?

Warning signs

  • ‘AI crawlers are allowed’, asserted from reading robots.txt alone.
  • A proprietary AI-readiness score with no log evidence behind it.
  • llms.txt presented as the main deliverable.
  • A recommendation to serve crawlers a special version of the page.
  • No distinction drawn between GPTBot and OAI-SearchBot.
  • Fixes reported as done on the day of deployment, with no crawler evidence.

Our page on choosing an AEO agency gives the broader vetting method, and AEO experts lists the skills to look for in the people doing the work.

Want proof of what AI crawlers receive?Send the domain and what sits in front of it. We request your key templates as each documented crawler and return the status codes, response sizes and first tickets.

Request the access review

Yes. Technical answer engine optimization, technical GEO, technical AI SEO and technical SEO for AI search describe one layer of work. The narrower phrases AI crawler access and AI crawler optimization name its first half. A buyer looking for a technical AEO agency and a buyer looking for help with an AI crawler robots.txt policy are asking for the service on this page.

The vocabulary is untangled on AEO vs GEO vs LLM SEO. The same practice appears under other labels on generative engine optimization and AI SEO agency, and the checklist view of this subject is our LLM SEO page.

Neighboring pages that a buyer scoping this layer tends to open next.

Find out what AI crawlers actually receive from your site

Send the domain and tell us what sits in front of it. We request your key templates as each documented crawler, read the logs you can share and return the first tickets in writing.

Request the access review

Frequently asked questions

What does technical AEO mean, in plain terms?
It is the configuration and code layer of answer engine optimization. It makes sure the documented AI crawlers are permitted in robots.txt, are not refused by a CDN or firewall, receive the page text in the HTML response, and find one canonical URL with markup that matches the visible page. Each fix is verified in server logs.
What does the technical side of AEO change on a website?
Usually a short list: robots.txt groups for named crawlers, CDN and firewall bot rules, IP allowlists, the rendering mode of templates whose content loads in the browser, canonical tags, redirects, sitemaps and structured data. The agency writes each change as a ticket with a test, implements it or pairs with your engineers, and files the evidence.
Which crawlers does OpenAI document, and what is each for?
Four. OAI-SearchBot surfaces sites in ChatGPT’s search features. GPTBot crawls content that may be used to train foundation models. ChatGPT-User visits pages for actions a user starts. OAI-AdsBot checks landing pages submitted as ads. OpenAI says the settings are independent, so a site can allow search while disallowing training.
What is the difference between ClaudeBot, Claude-User and Claude-SearchBot?
Anthropic describes three agents. ClaudeBot collects web content that could contribute to model training. Claude-SearchBot indexes content to improve search results for users. Claude-User retrieves a page when someone asks Claude a question that needs it. Each can be allowed or disallowed separately in robots.txt, on every subdomain.
If we disallow GPTBot, can ChatGPT search still show our pages?
Yes. OpenAI documents GPTBot as its training crawler and OAI-SearchBot as the agent behind ChatGPT’s search features, and says each setting is independent. Disallowing GPTBot signals that content should not be used for training. Disallowing OAI-SearchBot is what removes a site from ChatGPT search answers, apart from navigational links.
Do user-triggered fetchers such as ChatGPT-User and Perplexity-User obey robots.txt?
Not reliably, by the operators’ own description. OpenAI says robots.txt rules may not apply to ChatGPT-User because a user starts the request. Perplexity says Perplexity-User generally ignores robots.txt. Google says its user-triggered fetchers generally ignore it too. Anthropic is the exception: it says Claude-User can be disabled in robots.txt.
What happens if robots.txt returns a server error?
Under RFC 9309, a robots.txt that cannot be reached because of a server error means the crawler must assume a complete disallow. Google documents that it stops crawling for the first 12 hours and then uses the last good copy for up to 30 days. A file that fails during deployments is an access fault even if its contents are right.
Does AWS WAF block AI bots by default?
If the managed Bot Control rule group is enabled, yes. Its CategoryAI rule inspects for artificial intelligence bots, has a documented default action of Block, and applies to verified and unverified bots alike. Sites that turned Bot Control on to stop scrapers need that rule’s action overridden for the AI search crawlers they want to admit.
What do Cloudflare’s AI bot policies change for new domains?
Cloudflare’s documentation gives September 15, 2026 as the date from which new domains default to blocking bots it classifies as Training or Agent on pages that display ads, while Search bots stay allowed. Crawlers that mix search and training are caught by any option that blocks training. The policy is set per zone, so every domain is checked on its own.
How do you confirm a request really came from an AI crawler?
By checking the client IP address against what the operator publishes. OpenAI, Anthropic, Perplexity, Amazon, DuckDuckGo and Microsoft publish IP lists. Google, Apple and Common Crawl also support reverse DNS checks. A user agent string alone proves nothing, because anyone can send it, and Common Crawl warns that its own name is imitated.
Which vendors say their crawlers render pages, and which say nothing?
Google documents a rendering queue that executes JavaScript, and Apple says Applebot may render pages. The crawler documentation from OpenAI, Anthropic and Perplexity does not mention executing scripts. Microsoft advises against hiding answers in tabs because AI systems may not render them. We therefore test the raw HTML response for every template.
Is server-side rendering required for AI search visibility?
No vendor requires a particular rendering method. What matters is that the text is in the HTML the server returns, which server-side rendering, static generation and prerendering all achieve. Google itself recommends server-side or pre-rendering on the grounds that not all bots can run JavaScript. Fully client-rendered content is the risk.
Does Google require structured data for AI Overviews?
No. Google’s guidance says structured data is not required for its generative AI features, that no special schema.org markup exists for them, and that markup should match the visible text. It still recommends structured data for rich results. We ship a modest, accurate graph and treat disagreement between markup and page as a fault.
Does Google Search read llms.txt files?
No. Google states that Search does not use llms.txt or similar files, and that maintaining one neither helps nor harms visibility or rankings there. It adds that keeping one for other systems is fine. We add llms.txt when it can be generated from live content and kept accurate, never as a substitute for access or rendering work.
How long does a robots.txt change take to reach AI search systems?
It varies by operator. OpenAI says its search systems can take about 24 hours to adjust, and Amazon gives a similar figure. DuckDuckGo says a DuckAssistBot opt-out takes effect after 72 hours. RFC 9309 lets a crawler rely on a cached copy for up to 24 hours. We schedule verification after those intervals.
What log data do you need for a crawler access review?
Thirty to ninety days of request logs from the outermost layer, usually the CDN, plus the origin server’s access logs. The fields that matter are timestamp, host, path, status code, bytes sent, user agent and the client IP address as seen at the edge. Logs can be filtered to bot traffic before they are shared.
What if our host does not provide access logs?
Then the evidence comes from what the platform does expose: a CDN’s AI crawler report, firewall event logs, or a logging rule added to record requests from the documented user agents. If nothing can be captured, the review rests on synthetic requests and says so plainly, because a simulated request is weaker proof than a logged one.
When is a fix for AI crawler access considered verified?
When the operator’s own crawler, confirmed by its published IP range or reverse DNS, is logged receiving a 200 response of the expected size for the corrected page. A configuration diff and a passing synthetic request come first, but neither closes the ticket. Citations that follow are reported and never promised.
What is the price range for technical AEO work?
A review with tickets falls in the published AEO audit band of $1,000 to $4,000. Implementation is scoped afterward: configuration changes are small, while rebuilding rendering on a client-rendered site sits in the published project band of $25,000 to $100,000. Ongoing monitoring is part of a monthly retainer. A quote follows a written scope.
How long does the technical side of AEO take?
Plan on two to six weeks for one site from log access to closed tickets, if changes can be deployed as they are approved. Configuration fixes come first and are quick. Rendering changes take longer and are scoped separately. Approval and release schedules, not the work, usually set the pace.
Is technical work for AI search a one-time project or ongoing?
Both. The review and the fixes are a project with an end. Monitoring is ongoing, because access breaks quietly when a firewall rule, a framework version or a CDN changes. Scheduled requests per template and agent, a daily robots.txt comparison and a recurring log review keep the fixes in place.
Does technical work for answer engines replace technical SEO?
No. Google says its AI features rest on the same technical requirements as Search, so crawlability, indexing and canonical hygiene still carry the load. The AEO layer adds access for the other operators’ crawlers, tests that do not assume rendering, and verification in logs. The two are best run by one team against one ticket list.
What is Web Bot Auth?
Web Bot Auth is a method of signing HTTP requests cryptographically so that a website can verify which bot sent them against a public key directory. Cloudflare uses it to verify bots, AWS WAF labels requests by signature status, and Google says it is experimenting with it for its Google-Agent fetcher. It supplements IP lists and reverse DNS.
Does a Crawl-delay rule work for AI crawlers?
Only for agents whose operators support it. Crawl-delay is not part of the Robots Exclusion Protocol standard. Anthropic says its crawlers honor it, while Apple says Applebot does not follow it. For load problems, a page cache and a correct 429 response are more dependable than a directive that some crawlers ignore.

Want proof of what AI crawlers receive?Send the domain and what sits in front of it. We request your key templates as each documented crawler and return the status codes, response sizes and first tickets.

Request the access review

Get a free marketing proposal

Tell us what you are trying to grow and we will come back with a plan, not a pitch deck. Same-day reply on weekdays.

Privacy Preferences
When you visit our website, it may store information through your browser from specific services, usually in form of cookies. Here you can change your privacy preferences. Please note that blocking some types of cookies may impact your experience on our website and the services we offer.
Contact Us