Updated October 2026 · Written and maintained by the Progression Agency strategy team
Technical AEO is the engineering layer of answer engine optimization: the server, CDN, markup and URL work that lets the crawlers behind ChatGPT, Claude, Perplexity, Gemini, Microsoft Copilot and Google AI Overviews request a page, receive its full text and credit it to the right address. A technical AEO agency does that work as a service for companies whose sites sit behind a CDN or firewall, run on a JavaScript framework or span many templates and hosts, and it hands back each fix with proof from the server logs. Progression Agency is based in New York City and serves clients throughout the United States and worldwide.
On this page · 16 sections
- What is technical AEO?
- Which AI crawlers and user agents does the work cover?
- How does robots.txt apply to AI crawlers?
- Where do firewalls, CDNs and bot rules stop AI crawlers?
- Can AI crawlers read a page without JavaScript?
- What structured data belongs in the technical layer?
- How do canonical tags and URL hygiene fit in?
- Is llms.txt worth adding?
- How is crawler access verified in log files?
- How are the fixes verified?
- What does an engagement deliver?
- How long does the technical work take?
- How much does the technical layer cost?
- How do you choose a technical AEO agency?
- Is this the same as technical GEO or technical SEO for AI search?
- Related services
The short answerEvery major AI vendor documents its crawlers: OpenAI lists four user agents, Anthropic three and Perplexity two, while Google governs AI use through Googlebot and the Google-Extended token. The technical layer makes sure each agent you want is allowed in robots.txt, is not stopped by a firewall or bot rule, receives a 200 response with the text in the HTML, and finds one canonical URL per page with markup that matches what is visible. Each fix is written as a ticket with a pass-or-fail test, checked with a request that copies the crawler’s user agent, then confirmed in access logs against the vendor’s published IP ranges. An access fix is confirmed when the crawler next visits and is logged receiving the page. Whether an assistant then cites the page depends on content and corroboration, which the rest of an AEO program covers.
Agent names, stated purposes and controls are reported as each operator publishes them and were last checked on October 4, 2026. Operators add and rename agents, so the list is rechecked before every engagement. Passages marked editorial describe how we work. Ubersuggest’s United States data for October 2026 shows little or no search volume for this page’s phrases so far. Prices are the AEO planning bands this site publishes, and scoping in writing comes before any quote. No client is described and no result is claimed.
What is technical AEO?
It is the part of answer engine optimization that happens in configuration and code. It settles four things: whether the documented AI crawlers are permitted to request a page, whether anything between the internet and the server turns them away, whether the response contains the text, and whether the page declares one address and accurate markup for itself. Content, reputation and measurement are built on top of it and are described on our answer engine optimization agency page.
It is sold as a service, not handed over as a checklist, because the faults live in systems marketing teams rarely control: a CDN account, a web application firewall, a framework’s rendering mode, a deployment pipeline. The job is to find the fault, write the change so an engineer can ship it, and prove afterward that the crawler now receives the page.
How does it differ from technical SEO?
Technical SEO serves search engines that render JavaScript, publish detailed webmaster guidelines and report crawl problems in a console. The AEO layer serves a longer list of crawlers, most of which document only their names, purposes and IP addresses. The foundations are shared, and Google states that its AI features have no technical requirements beyond those of Search. The added work is access per vendor, tests that do not assume rendering, and verification in logs because there is no console to ask. Crawl budget, Core Web Vitals, migrations and index management stay with our technical SEO agency team.
How does it differ from the pages next to it?
Several pages on this site touch the technical side of AI visibility. Each has one job, and this page is the service that implements and proves the fixes.
| Topic | This page | Covered in full on |
|---|---|---|
| The item-by-item review a developer can follow | Summary only | LLM SEO: the technical side |
| How retrieval and citation work inside an assistant | No | How AI search works |
| Whether URL length, depth or wording matter | No; only canonical and redirect fixes as service items | URLs and AI citation |
| The audit as a fixed-price product | No | AEO audit |
| WordPress sources of robots rules, plugins and caching | No | AEO for WordPress |
| Approval paths and release governance in large companies | No | Enterprise AEO |
| Crawler documentation by vendor, edge and firewall rules, rendering by framework, log verification and proof of fixes | Yes | This page |
Who needs this as a service?
- Sites behind Cloudflare, AWS WAF, Vercel or a host firewall whose bot settings nobody has reviewed since AI crawlers appeared.
- Applications built with React, Vue or Angular where part of the content is assembled in the browser.
- Sites with many templates, hosts or subdomains, where one robots.txt or one rule cannot be assumed to cover everything.
- Teams that have been told ‘AI crawlers are allowed’ and want that shown in logs.
- Companies about to change platform or CDN that want crawler access tested before and after the move.
- Publishers and software companies that want to admit search crawlers while declining training crawlers, implemented exactly.
Which AI crawlers and user agents does the work cover?
Every agent an operator documents, grouped by what it is for. The operators publish three kinds: crawlers that build a search index, crawlers that collect training data, and fetchers that retrieve one page because a user asked. Each kind is controlled separately, and that separation is what makes a precise policy possible. The documentation to configure against is published by OpenAI, Anthropic, Perplexity and Google.
| Operator | Agent or token | Stated purpose | robots.txt behavior stated | Identity check published |
|---|---|---|---|---|
| OpenAI | OAI-SearchBot | Surfaces sites in ChatGPT’s search features | Honored; opting out removes the site from ChatGPT search answers | IP range list |
| OpenAI | GPTBot | Crawls content that may be used to train foundation models | Honored | IP range list |
| OpenAI | ChatGPT-User | Visits a page for a user action in ChatGPT or a custom GPT | Rules may not apply, because a user started the request | IP range list |
| OpenAI | OAI-AdsBot | Checks landing pages submitted as ChatGPT ads | Visits only pages submitted as ads | IP range list |
| Anthropic | ClaudeBot | Collects web content that could contribute to model training | Honored, including Crawl-delay | Published IP list |
| Anthropic | Claude-SearchBot | Indexes content to improve search results | Honored | Published IP list |
| Anthropic | Claude-User | Retrieves a page when a Claude user asks | Can be disabled in robots.txt | Published IP list |
| Perplexity | PerplexityBot | Surfaces and links sites in Perplexity search; not used for training | Controlled in robots.txt | IP range list |
| Perplexity | Perplexity-User | Fetches a page to answer a user’s question | Generally ignores robots.txt | IP range list |
| Googlebot | Crawls for Search, which AI Overviews and AI Mode are built on | Honored | Reverse DNS and IP range lists | |
| Google-Extended | Token only: governs use for Gemini training and grounding | A robots.txt token, not a crawler; no effect on Search | Not applicable | |
| Google-Agent and other user-triggered fetchers | Act on a user’s request | Generally ignore robots.txt | IP range lists; Web Bot Auth in testing | |
| Microsoft | Bingbot | Bing’s crawler; Microsoft ties Copilot citations to Bing’s index and controls | Honored, with robots meta controls | IP range list |
| Apple | Applebot | Search in Siri, Spotlight and Safari | Honored; follows Googlebot rules if not named; ignores crawl-delay | Reverse DNS and IP range list |
| Apple | Applebot-Extended | Token only: opts content out of training Apple’s foundation models | A robots.txt token | Not applicable |
| Amazon | Amazonbot, Amzn-SearchBot, Amzn-User | Service improvement and possible training; search such as Alexa; user actions | First two honored; Amzn-User may not follow every directive | Published IP addresses |
| DuckDuckGo | DuckAssistBot | Crawls in real time for AI-assisted answers; not used for training | Honored; an opt-out takes effect after 72 hours | Published IP addresses |
| Mistral | MistralAI-User, MistralAI-Index | User actions; indexing for Mistral search; neither used for training | Controlled with robots.txt tags | Published IP addresses |
| Common Crawl | CCBot | Builds an open repository of web crawl data | Honored | Reverse DNS and IP range list |
Three kinds of request: search, training and user-triggered
Search crawlers decide whether a site can be found and linked in an assistant’s answers: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot. Training crawlers and tokens govern model training: GPTBot, ClaudeBot, Google-Extended and Applebot-Extended. User-triggered fetchers retrieve one page because a person pasted a link or asked a question: ChatGPT-User, Claude-User and Perplexity-User. OpenAI states that each of its settings is independent of the others, so a site can admit search and decline training. The work starts by writing that decision down, because a rule aimed at training that also catches a search crawler takes the site out of answers.
Tokens that never appear in a log
Google-Extended and Applebot-Extended are not crawlers. Google describes Google-Extended as a standalone product token with no user agent string of its own: the crawling is done by Google’s existing agents, and the token controls whether content may be used to train Gemini models and to ground answers in Gemini Apps and Vertex AI. Google adds that the token does not affect inclusion or ranking in Search. Apple’s Applebot-Extended works the same way for Apple’s foundation models. A log review that waits for either name to show up is waiting for something that cannot happen.
How long does a change take to register?
Not instantly. OpenAI says its search systems can take about 24 hours to adjust after a robots.txt update, and Amazon gives the same figure for its agents. DuckDuckGo says an opt-out for DuckAssistBot takes effect after 72 hours. The protocol itself lets a crawler reuse a cached robots.txt for up to 24 hours. Google says excluding a site with its Search Console generative AI control generally takes a few days. Verification is scheduled around those intervals, so that a correct fix is not reported as a failure on day one.
Want proof of what AI crawlers receive?Send the domain and what sits in front of it. We request your key templates as each documented crawler and return the status codes, response sizes and first tickets.
How does robots.txt apply to AI crawlers?
By the same standard as for any crawler, the Robots Exclusion Protocol published as RFC 9309. The parts that matter for AI access are how a crawler picks its group, what happens when the file cannot be fetched, and who else is writing to the file.
Which group does a crawler obey?
A crawler looks for the group whose user-agent line matches its product token, compared without regard to case, and obeys that group alone. The wildcard group applies only when no named group matches, and within a group the most specific matching path wins. Two consequences follow. A short group that names OAI-SearchBot and allows everything is not cancelled by a strict wildcard group further down. And a group that names a crawler in order to block one folder replaces the wildcard rules for that crawler, so anything the wildcard group disallowed is open to it again unless those lines are repeated.
What happens when robots.txt itself fails?
The standard separates two failures. If the file is unavailable, which for HTTP means a status in the 400 range, a crawler may access any resource. If it is unreachable because of a server or network error, a status in the 500 range, the crawler must assume everything is disallowed. Google documents its own handling: on a server error it stops crawling the site for the first 12 hours, then falls back on the last good copy of the file for up to 30 days. A robots.txt that returns a 500 or 503 for a few hours during each deployment is therefore an access fault, even though the contents of the file are correct. We test the file’s status from outside, repeatedly, including during a release.
Who else is writing the file?
The file a crawler receives is not always the file in the repository. Cloudflare’s managed robots.txt setting, when switched on, places Cloudflare’s own block ahead of the origin’s file, with disallow groups for named crawlers that include GPTBot, ClaudeBot, Google-Extended, CCBot and Amazonbot. Shopify generates the file itself and lets a theme template override it. Frameworks and CMS plugins produce theirs at build time or on request. The check is always made against the public URL of each host, and the result is compared with what the team believes it published. WordPress has its own set of sources, listed on AEO for WordPress.
Does Crawl-delay work?
For some agents. Crawl-delay is not part of RFC 9309. Anthropic says its crawlers support it; Apple says Applebot does not follow it. Where server load is the worry, Google’s guidance is not to answer with 401 or 403 as a way to slow crawling, and to treat 429 as the signal for too many requests. A page cache in front of the origin is the ordinary remedy for load, and it turns no one away.
| Goal | Allow | Disallow | Note |
|---|---|---|---|
| Be cited by assistants and permit training | Every documented agent | Private paths only | The simplest file; each agent is still tested |
| Be cited, decline training | OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot and the user fetchers | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot | The operators document these settings as independent |
| Stay in Google and Bing only | Googlebot, Bingbot | The named AI search crawlers | Per OpenAI, the site then leaves ChatGPT search answers, though it may still appear as a navigational link |
| Keep one section out of answers | Everything else | That path, in every named group | Named groups do not inherit wildcard rules, so the path is repeated |
| Reduce load without blocking | All agents | Nothing | Use caching, and Crawl-delay where it is supported; avoid 403 |
Where do firewalls, CDNs and bot rules stop AI crawlers?
In front of the server, where robots.txt has no say. A robots.txt rule is a request that well-behaved crawlers honor; a firewall rule is enforcement. The operators expect this layer to matter: OpenAI recommends allowing its published IP ranges as well as its robots.txt token, Perplexity publishes firewall instructions for Cloudflare and AWS, and Anthropic says its crawlers will not try to get past a CAPTCHA.
Cloudflare
Cloudflare sorts AI traffic into Search, Agent and Training in its AI bot policies and lets each be allowed, blocked everywhere or blocked only on pages that carry ads. Its documentation gives September 15, 2026 as the date from which new domains default to blocking Training and Agent bots on pages with ads while leaving Search allowed, and notes that a crawler used for both search and training is caught by any option that blocks training. Bot Fight Mode issues computational challenges to traffic it identifies as bots and cannot be adjusted with custom firewall rules. AI Crawl Control reports which AI services request a site and sets allow or block rules per crawler. We read all three settings, zone by zone, before touching robots.txt.
AWS WAF
The AWS managed Bot Control rule group contains a rule named CategoryAI that inspects for artificial intelligence bots. Its documented default action is Block, and AWS notes that the action applies to verified and unverified bots alike. A site that enabled Bot Control against scrapers is therefore refusing the documented AI search crawlers unless that rule’s action was overridden. AWS also documents that Bot Control verifies bots by the origin IP address of the request, which matters when a proxy or load balancer sits in front and the true client address travels in a forwarded header.
Vercel and other platform firewalls
Vercel offers an AI bots managed ruleset that identifies known AI crawlers and can be set to log or deny, and a separate bot protection ruleset that serves a JavaScript challenge to non-browser clients while exempting verified bots. Other hosts and security products have equivalents under different names. The method does not change with the product: find every rule that matches on user agent, bot category, country or request rate, then test what each documented crawler receives.
Challenge pages, rate limits and country blocks
A challenge page is HTML that is not the content, often delivered with a 200 or 403 status. To a crawler that does not solve challenges, it is the page. Rate limits answer fast crawlers with 429 or 503. Country blocks stop crawlers whose IP ranges sit outside the permitted countries. Each of these shows in logs as a status or response size that differs from what a browser receives for the same URL, which is why verification relies on logs and not on screenshots.
| Layer | Setting or rule | Documented behavior | What the service does |
|---|---|---|---|
| CDN bot policy | Cloudflare AI bot policies | Search, Agent and Training are each allowed or blocked; new-domain defaults block Training and Agent on pages with ads | Sets each zone to match the written policy |
| CDN bot challenge | Cloudflare Bot Fight Mode | Challenges identified bots; not adjustable with custom rules | Agrees on or off with the security owner; tests each crawler |
| Managed firewall rules | AWS Bot Control, CategoryAI | Default action Block, for verified and unverified bots | Overrides the action or narrows the rule |
| Platform firewall | Vercel AI bots ruleset | Log or deny for known AI bots | Sets it to log unless the policy is to block |
| IP allowlists | Operator IP range lists | OpenAI and Perplexity recommend admitting their published ranges | Automates updates from the published lists |
| robots.txt at the edge | Cloudflare managed robots.txt | Adds disallow groups for named crawlers ahead of the origin file | Turns it off or reconciles it with the origin file |
For a comparison of those two platforms themselves, see Cloudflare vs AWS.
Can AI crawlers read a page without JavaScript?
Assume they cannot, and build so that it does not matter. Google renders JavaScript. Most other operators say nothing on the subject, and Microsoft warns that AI systems may not render hidden content. The engineering target is that the text an assistant would quote is present in the HTML the server sends.
What do the operators say about rendering?
Google’s JavaScript documentation describes a rendering queue in which a headless Chromium executes scripts after the first fetch, and adds that server-side rendering or pre-rendering is still a good idea because not all bots can run JavaScript. Its note on dynamic rendering calls that technique a workaround, not a recommendation, and observes that other search engines may ignore JavaScript. The crawler documentation from OpenAI, Anthropic and Perplexity describes what each agent is for and how to control it; none of it states that the agent executes scripts. Apple says Applebot may render pages and advises that sites degrade gracefully when resources cannot be loaded. Microsoft’s guidance for AI search answers advises against hiding important answers in tabs or expandable menus, because AI systems may not render them, and against keeping key information only in images or PDFs.
What do the frameworks do by default?
It depends on the framework and, within a framework, on choices made per route. The table gives each one’s documented default and the place to look for content that has left the server response.
| Framework | Documented default | What to check |
|---|---|---|
| Next.js (App Router) | Layouts and pages are Server Components | Client Components that fetch the main content after load |
| React without a framework | Server APIs can render components to HTML; a framework normally calls them | Whether any server rendering is configured at all |
| Angular | Applications are client-side rendered; server-side and hybrid rendering are opt-in | Whether SSR or prerendering is switched on for public routes |
| Vue | Components render in the browser; server rendering and static generation are available | Pages that open on a loading state and then fetch their content |
| SvelteKit | Pages render on the server first, then hydrate; ssr, csr and prerender are page options | Routes where server rendering has been turned off |
| Astro | Mostly static HTML, with interactive islands | Content placed inside client-only islands |
Our Next.js development and React development teams implement the change where a client has no engineers to spare. Headless WordPress and headless CMS vs traditional CMS cover decoupled front ends, and hosting for React and Next.js covers where server rendering runs.
The raw-HTML test
The test is small, repeatable and independent of any tool’s score. It is run per template, not per page.
- List the templates that carry commercial content: home, service or product, pricing, location, article, FAQ and comparison.
- For each template, choose one URL and three sentences an assistant would need in order to answer a buyer.
- Request the URL with a command-line client that runs no scripts: once with a browser user agent, once with each documented crawler’s.
- Search each response body for the three sentences.
- Record the status code, the response size and whether each sentence was found.
- Where a sentence is missing, identify the component that loads it and the data source behind it.
- Write the fix as server rendering, prerendering or moving the text into the first response, with the same three sentences as its test.
What usually goes missing?
- Pricing tables filled from an API call after the page loads.
- Product specifications inside tabs that fetch their panel on click.
- Reviews and ratings injected by a third-party script.
- FAQ answers revealed by a script instead of sitting in the markup.
- Store, dealer and location finders that exist only as a map.
- Content held back by a cookie or consent layer until someone clicks.
- Text inside iframes, images and PDFs.
- Views that exist only as URL fragments, with no address of their own.
Running React, Vue or Angular?Tell us the framework and hosting. We run the raw-HTML test on each template and show which content never reaches the server response.
What structured data belongs in the technical layer?
Enough to state plainly who publishes the page and what it describes, and nothing the page does not show. Structured data is confirmation for a machine. It is not a route into answers.
What do Google and Microsoft each say?
Google’s guide to generative AI features in Search says structured data is not required for them, that no special schema.org markup exists for AI, and that markup remains worth using for rich results provided it matches the visible text. Its structured data policies rule out marking up content readers cannot see and markup that misrepresents the page. Microsoft’s guidance says schema helps search engines and AI systems understand content, and recommends the JSON-LD format. The two positions fit together: accurate markup helps identification, and neither company treats it as a substitute for the text.
The graph we ship
- Organization, with legal name, logo, address, contact point and links to verified profiles.
- WebSite and WebPage, tying each URL to the site and to its breadcrumb.
- Service, or Product with Offer, where the page describes something sold and states a price.
- Article with author and dates on editorial pages, where dates change only when the content does.
- FAQPage where the page shows real questions with their full answers.
- BreadcrumbList on every page below the home page.
All of it is JSON-LD in schema.org vocabulary, one connected graph per page, generated from the same fields that produce the visible copy so that the two cannot drift apart.
Markup that contradicts the page
The fault we look for is disagreement: a price in the markup that differs from the price in the table, opening hours changed on the page but not in the plugin, a rating with no visible reviews, two plugins describing one page as two organizations. Each template is validated and then compared field by field with the rendered text. Our schema and copy validator automates the comparison, and the schema markup generator writes clean markup for hand-built sections.
How do canonical tags and URL hygiene fit in?
They decide which address gets the credit. Google’s generative AI performance report assigns most of its data to a page’s canonical URL, and Bing reports citations per URL, so a page reachable at four addresses splits its own evidence four ways.
One address per page
Google treats redirects and the rel=canonical link as strong canonical signals and sitemap inclusion as a weak one; the link relation itself is defined in RFC 6596. The service items are concrete: one protocol and host, one trailing-slash rule, parameters that create no indexable copies, redirects that reach the final address in one hop, and a sitemap that lists only canonical URLs with honest modification dates. Where the platform supports IndexNow, which Microsoft recommends for keeping AI answers current, changed URLs are submitted as they change.
Canonical in the HTML, not added by a script
Google’s JavaScript documentation says the best place to set a canonical is the HTML, and warns against using a script to change it to a value different from the one the server sent. A crawler that runs no scripts sees only the server’s version, so a canonical injected or rewritten in the browser is, for that crawler, missing or wrong. The same document advises routing single-page applications with the History API instead of URL fragments, so that each view has a real address.
- Redirect chains shortened to a single hop.
- http and https, www and bare-host variants resolved to one.
- A canonical tag present in the server response on every template.
- Parameterized duplicates canonicalized or kept out of the crawlable set.
- Error pages that answer 200 changed to return the correct status.
- Sitemap entries limited to canonical, indexable URLs that return 200.
Whether URL length, depth or wording matter is a separate question, answered on URLs and AI citation.
Is llms.txt worth adding?
It is an optional extra, added when it is cheap to generate and easy to keep accurate. It is not an access control, and it is not a ranking signal.
What does the proposal specify?
The llms.txt proposal describes a Markdown file, at the site root or at any path, that gives a short description of the site or section and links to the pages an agent would need, ideally to Markdown versions of them. Its current version also suggests link relations, in the HTML or in HTTP headers, that point from a page to its Markdown version and to the llms.txt file covering it. The format is deliberately plain: a title, a short summary and lists of links.
Who reads it?
Google states that Search does not use llms.txt, and that keeping one neither helps nor harms visibility there. The proposal’s own account is that the files are used most heavily for software documentation, where coding agents follow them to find references. The developer documentation sites of OpenAI and Perplexity point readers to an llms.txt index and to Markdown versions of their pages, which shows the convention at work for documentation. None of the crawler documentation we read says a search or training crawler gives the file special treatment.
When do we add one?
When the site has documentation, policies or product data that agents are likely to be sent to; when the CMS can generate the file from live content, so it never goes stale; and when every link in it resolves to a page that is indexable and current. We do not add one in place of fixing access or rendering. The free llms.txt generator produces a first draft.
How is crawler access verified in log files?
By finding the crawler’s own requests in the access logs and checking three things: that the request really came from the operator, what status and size the server returned, and which URLs were asked for. Logs are the only record of what a crawler was served. Every other test is a simulation.
Which logs, and which fields?
The log has to come from the outermost layer that answers requests. If a CDN sits in front, the origin’s log shows only what the CDN let through, and a crawler blocked at the edge never appears in it. We ask for 30 to 90 days from each layer: the CDN’s request logs at the edge, and the web server’s access log at the origin, whether that is Apache, NGINX or IIS. The fields needed are the timestamp, the requested host and path, the status code, the bytes sent, the user agent and the client IP address as seen at the edge.
How are real crawlers told from imitators?
A user agent string is a claim, not proof. Common Crawl itself warns that other crawlers falsely identify as CCBot. The operators publish stronger identifiers, and the review uses whichever one each offers.
| Operator | Method published | How the review uses it |
|---|---|---|
| OpenAI | A list of IP ranges for each agent | Match the logged address against the list for the agent named |
| Anthropic | A published list of source IP addresses | Match the logged address against the list |
| Perplexity | A list of IP ranges for each agent, with firewall instructions | Match the address; mirror the allow rule it describes |
| Reverse DNS to googlebot.com or google.com hosts, plus IP range lists per crawler class | Reverse lookup, then forward lookup, or match the lists | |
| Apple | Reverse DNS in the applebot.apple.com domain, plus a list of IP ranges | Reverse lookup or match the list |
| Microsoft | A list of Bingbot IP ranges | Match the logged address against the list |
| Amazon and DuckDuckGo | Published IP addresses for each agent | Match the logged address against the list |
| Common Crawl | Reverse DNS to crawl.commoncrawl.org, plus a list of IP ranges | Reverse lookup or match the list |
A newer method is cryptographic. Web Bot Auth signs HTTP requests so that the receiver can verify the sender against a public key directory. Cloudflare uses it as one way to verify bots, AWS WAF labels requests by signature status, and Google says it is experimenting with the protocol for its Google-Agent fetcher. Where a client’s edge supports it, signature status becomes one more column in the review.
What should the log show after a fix?
- Requests from the operator’s verified IP ranges carrying the expected user agent.
- A 200 status on the URLs that matter, and 301 only where a redirect is intended.
- Response sizes in line with what a browser receives for the same URL.
- A robots.txt fetch from each crawler that returns 200.
- No run of 403, 429 or 503 responses to a documented agent.
- Requests that reach deep templates, not only the home page.
What if there are no logs?
Some platforms do not expose raw access logs. The evidence then comes from whatever the platform does provide: a CDN’s AI crawler report, a firewall event log, or a logging rule added for the purpose that records requests from the documented user agents. Where nothing can be captured, the report says so and rests on synthetic tests alone, labeled as such.
- Collect logs from every layer and every host for the same period.
- Filter for the documented user agents.
- Verify each matching request against the operator’s IP list or by reverse DNS, and set imitators aside.
- Tabulate by agent, host, status code and template.
- Compare response sizes with a browser’s for the same URLs.
- List the URLs requested and the URLs never requested.
- Read the robots.txt fetches separately.
- Write each anomaly as a finding, with the log lines attached.
Need the ranges before a call?AEO audits, retainers and implementation projects are published as planning ranges. Ask for the band that fits your hosts and templates.
How are the fixes verified?
In four levels of increasing strength, and a fix is not closed until it reaches the third. A recommendation is not a result, and a deployment is not a result either until the crawler has been served the corrected page.
Tickets with acceptance tests
Each finding becomes a ticket an engineer can act on without a meeting: the fault, the evidence, the change, the systems touched, the risk, and a test that passes or fails. ‘Allow AI crawlers’ is not a ticket. A usable one reads like this: OAI-SearchBot receives 403 on the pricing template from the edge; override the Bot Control CategoryAI action for requests from OpenAI’s published ranges; test: a request with the OAI-SearchBot user agent returns 200 and contains the three reference sentences.
| Fix | Acceptance test | Evidence filed | Rechecked |
|---|---|---|---|
| robots.txt group added or corrected | The public file shows the group on every host; a parser test returns allowed for the target URLs | Before and after copies of the file, with timestamps | After each deploy |
| Firewall or bot rule changed | A request with the crawler’s user agent returns 200 and the full body | The rule diff; response headers and size | Weekly, and after any rule change |
| IP allowlist added | Requests from the operator’s ranges are not challenged | Edge log lines from verified addresses | When the operator’s list changes |
| Template server-rendered | The three reference sentences are present in the raw HTML | Saved response bodies | After each release |
| Canonical corrected | The server response carries one self-referencing canonical | A capture of the head for each template | After each release |
| Structured data reconciled | The validator passes and values equal the visible text | Validator output and a comparison table | Monthly |
| Redirect chain shortened | One hop to a 200 | Trace output | After URL changes |
| robots.txt availability | 200 on repeated fetches, including during a deploy | An uptime record for the file | Continuously |
Four levels of evidence
- Configuration: the rule, file or template has changed, shown by a diff.
- Synthetic request: a request copying the crawler’s user agent now receives the right status and body. This proves the path is open for that string, not that the operator’s network is admitted.
- Verified crawler request: the operator’s own crawler, confirmed by IP range or reverse DNS, is logged receiving a 200 with the full response.
- Downstream signal: the page shows as a cited URL in Bing Webmaster Tools, gains impressions in Search Console’s generative AI report, or is linked in a prompt-set answer.
Level three closes the ticket. Level four is reported when it happens and is never promised, because citation depends on content and competition as well as access.
Regression checks after every release
Access breaks quietly: a new firewall rule, a framework upgrade that changes a rendering default, a CDN migration, a plugin that rewrites robots.txt. The service therefore leaves monitoring behind it: scheduled requests per template and per documented user agent, an alert when a status code or response size changes, a daily copy of each host’s robots.txt with an alert on any difference, and a recurring log review. Where the client has a deployment pipeline, the raw-HTML test for the reference sentences is added to it, so that a release which removes the text fails before it ships.
What does an engagement deliver?
Shipped changes, each with its proof, and the monitoring that keeps them in place.
Deliverables
- A written access policy: which agents are admitted for search, for training and for user-triggered fetches.
- A crawler access report per host, built from logs, with verified requests separated from imitators.
- A layer map of every CDN, firewall, platform and application rule that touches bot traffic.
- A raw-HTML report per template, with the reference sentences found or missing.
- Tickets with acceptance tests, in priority order.
- Implementation where access allows, or pairing with your engineers where it does not.
- A reconciled structured data graph per template.
- Canonical, redirect and sitemap corrections.
- An evidence file for each closed ticket.
- Monitoring and a recurring log review.
What is not included
- Writing or restructuring page copy, which is AEO content writing.
- Prompt-set measurement and reporting, which belong to the AI visibility audit.
- Getting around an operator’s controls, or disguising requests.
- Serving crawlers different content from the content users receive.
- Guarantees of citation.
- Security decisions: we recommend, and your security owner approves.
How long does the technical work take?
Plan on two to six weeks for a single site, from access to the logs to closed tickets, when changes can be deployed as they are approved. That is a planning range. The work itself is small; elapsed time is governed by log access, change approval and release schedules.
| Phase | When | Work | Output |
|---|---|---|---|
| Access and inventory | Week 1 | Logs, CDN and firewall access; host and template inventory; written policy | Layer map and access policy |
| Review | Weeks 1 to 2 | Log analysis, requests per agent, raw-HTML tests | Findings with evidence |
| Tickets | Week 2 | Changes written with acceptance tests and priority | Ticket list |
| Implementation | Weeks 2 to 5 | Configuration changes first; rendering and template changes after | Deployed fixes |
| Verification | From each deploy, plus the operator’s stated delay | Synthetic tests at once; verified crawler requests as they arrive | Evidence files |
| Monitoring | Ongoing | Scheduled tests, robots.txt watch, recurring log review | Alerts and a short monthly note |
Rendering changes on a client-rendered application can take longer than everything else combined, and are scoped separately after the raw-HTML test. Approval time in large organizations is a subject of its own, covered on enterprise AEO. Citation timing after access is fixed is a different question; how long AEO takes deals with it.
How much does the technical layer cost?
It is priced from the published AEO planning ranges. A review with tickets falls in the audit band; implementation is scoped after the review, because rendering work varies so widely. A quote follows a written scope.
| Planning band | Engagement type | Technical work it holds |
|---|---|---|
| $1,000 to $4,000, one time | AEO audit | The review: logs, access per agent, raw-HTML tests and a ticket list (the band is published for 10 to 20 hours) |
| $25,000 to $100,000, per project | Implementation | A full rendering and restructuring program on a client-rendered site |
| $1,500 to $5,000 per month | Retainer, small business | Monitoring and fixes as part of a wider AEO program |
| $5,000 to $10,000 per month | Retainer, mid-market | The same, across more templates and platforms |
| $10,000 to $20,000 or more per month | Retainer, enterprise | The same, across many hosts and teams |
What moves a technical quote?
- Rendering: a server-rendered site needs configuration; a client-rendered one may need engineering.
- The number of templates, hosts and subdomains.
- The number of layers in front of the origin, and who controls each.
- Whether raw logs exist, and how far back they go.
- Who implements: your engineers with our tickets, or our developers.
- Whether monitoring continues inside a retainer.
The full set of ranges and what drives them is on AEO pricing and in the marketing agency pricing guide.
How do you choose a technical AEO agency?
Choose on evidence habits. The work is verifiable end to end, so a provider’s method shows in what they ask for on the first day and what they hand over on the last.
| What to require | The check |
|---|---|
| Works from logs | They ask for edge and origin logs before proposing fixes |
| Names crawlers precisely | They separate search, training and user-triggered agents, operator by operator |
| Knows the operators’ documentation | Ask what each operator says its agent is for and how it is controlled |
| Tests without rendering | They show raw HTML responses, not browser screenshots |
| Verifies identity | They check IP ranges or reverse DNS, not user agent strings alone |
| Writes testable tickets | Each ticket carries a pass-or-fail test |
| Proves fixes | They return log lines from verified crawlers after deployment |
| Leaves monitoring behind | Scheduled tests and a robots.txt watch remain when they finish |
| Respects security ownership | Firewall changes are proposed to your security team, not made around it |
What should you ask before signing?
- Which crawlers will you test, and what does each one’s operator say it is for?
- Which logs do you need, from which layers, and covering how long?
- How will you tell a real crawler from a request that copies its user agent?
- What do you do when our platform exposes no logs?
- How will you test rendering without a browser?
- What does one of your tickets look like?
- What evidence closes a ticket?
- Who makes the change in the CDN, the firewall and the application, and who approves it?
- What monitoring stays in place, and who receives the alerts?
- Which parts of this are one-time, and which need a retainer?
Warning signs
- ‘AI crawlers are allowed’, asserted from reading robots.txt alone.
- A proprietary AI-readiness score with no log evidence behind it.
- llms.txt presented as the main deliverable.
- A recommendation to serve crawlers a special version of the page.
- No distinction drawn between GPTBot and OAI-SearchBot.
- Fixes reported as done on the day of deployment, with no crawler evidence.
Our page on choosing an AEO agency gives the broader vetting method, and AEO experts lists the skills to look for in the people doing the work.
Want proof of what AI crawlers receive?Send the domain and what sits in front of it. We request your key templates as each documented crawler and return the status codes, response sizes and first tickets.
Is this the same as technical GEO or technical SEO for AI search?
Yes. Technical answer engine optimization, technical GEO, technical AI SEO and technical SEO for AI search describe one layer of work. The narrower phrases AI crawler access and AI crawler optimization name its first half. A buyer looking for a technical AEO agency and a buyer looking for help with an AI crawler robots.txt policy are asking for the service on this page.
The vocabulary is untangled on AEO vs GEO vs LLM SEO. The same practice appears under other labels on generative engine optimization and AI SEO agency, and the checklist view of this subject is our LLM SEO page.
Related services
Neighboring pages that a buyer scoping this layer tends to open next.
- Technical SEO agency: crawling, indexing, speed and migrations for search engines.
- Answer engine optimization agency: the full program this layer supports.
- LLM SEO: the technical checks, written for the person implementing them.
- How AI search works: the retrieval pipeline behind an answer.
- URLs and AI citation: which URL issues matter and which do not.
- AEO audit: the fixed-scope review and what it should contain.
- AEO for WordPress: the same layer on WordPress specifically.
- AEO services: every deliverable in an AEO program, with its evidence.
- Enterprise AEO: approvals and ownership at scale.
- AI crawler access checker: a free first pass on robots.txt and status codes.
- Extractable content checker: a free look at what a page yields without scripts.
- What is AEO?, AEO vs SEO and AEO tools: background for colleagues new to the subject.
- AEO pros and cons and AEO myths: what to weigh, and what to disbelieve, before a purchase.
- Local AEO: where listings matter more than crawling.
- Website speed optimization and website migration services: the neighboring engineering projects.
- SEO audit service: the ranking-side review of the same site.
Find out what AI crawlers actually receive from your site
Send the domain and tell us what sits in front of it. We request your key templates as each documented crawler, read the logs you can share and return the first tickets in writing.
Getting found in search
AI, AEO and what is changing
Paid media and lead generation
Websites and design
Choosing and working with an agency
Software and app development
Website design by industry and type
Web development, platforms and hosting
Social, content and brand
By industry and by situation
Frequently asked questions
What does technical AEO mean, in plain terms?
What does the technical side of AEO change on a website?
Which crawlers does OpenAI document, and what is each for?
What is the difference between ClaudeBot, Claude-User and Claude-SearchBot?
If we disallow GPTBot, can ChatGPT search still show our pages?
Do user-triggered fetchers such as ChatGPT-User and Perplexity-User obey robots.txt?
What happens if robots.txt returns a server error?
Does AWS WAF block AI bots by default?
What do Cloudflare’s AI bot policies change for new domains?
How do you confirm a request really came from an AI crawler?
Which vendors say their crawlers render pages, and which say nothing?
Is server-side rendering required for AI search visibility?
Does Google require structured data for AI Overviews?
Does Google Search read llms.txt files?
How long does a robots.txt change take to reach AI search systems?
What log data do you need for a crawler access review?
What if our host does not provide access logs?
When is a fix for AI crawler access considered verified?
What is the price range for technical AEO work?
How long does the technical side of AEO take?
Is technical work for AI search a one-time project or ongoing?
Does technical work for answer engines replace technical SEO?
What is Web Bot Auth?
Does a Crawl-delay rule work for AI crawlers?
Want proof of what AI crawlers receive?Send the domain and what sits in front of it. We request your key templates as each documented crawler and return the status codes, response sizes and first tickets.
Get a free marketing proposal
Tell us what you are trying to grow and we will come back with a plan, not a pitch deck. Same-day reply on weekdays.
