Migrations break crawler access in predictable ways
A site migration, whether a domain change, platform move, or infrastructure switch, risks breaking access for every crawler identity simultaneously, including AI crawlers that standard migration checklists do not always name explicitly. Google’s own site-move guidance covers redirects, canonical consistency, and verification timing for its own crawler; the same mechanics apply to AI crawlers, with a few additions worth checking separately.
Before the migration
Capture a baseline: current robots.txt rules per crawler identity, current sitemap contents, and a sample of pages’ delivered HTML and structured data. Without a baseline, you cannot confirm the migration preserved behavior rather than silently changing it.
During cutover
Confirm the new infrastructure serves the same robots.txt rules for every AI crawler identity that mattered on the old site; a new CDN or WAF default configuration can silently introduce a block that was never present before. Confirm 301 redirects (not soft redirects or client-side JavaScript redirects) map old URLs to new ones, since some AI crawlers may not execute JavaScript the way a browser does. Confirm the new sitemap reflects the new URL structure and is reachable at its expected location.
After the migration
Re-run the same page sample checked in the baseline and diff the results: robots directives, sitemap presence, structured data, and delivered content. Watch server logs for AI-crawler-identified traffic hitting old URLs; a sustained pattern of 404s on old paths past the expected redirect window indicates a missed redirect rule, not just stale caching.
Common migration mistakes specific to AI crawlers
- Assuming a redirect rule written for browsers also applies to non-browser clients if it relies on JavaScript.
- Forgetting to update crawler-specific robots.txt groups when copying the file to new infrastructure.
- Losing structured data during a platform or template change without re-auditing it.
- Declaring the migration complete based only on human-visible browsing, without checking raw-fetch results.
Keep a rollback plan specific to crawler access
If post-migration monitoring shows an AI crawler identity losing access it previously had, have a defined rollback or hotfix path for the specific layer that changed: DNS, CDN configuration, robots.txt, or redirect rules. Diagnosing which layer caused the regression before rolling back anything blindly saves time, since a migration typically changes several layers at once and only one is usually the actual cause.
See website migration seo checklist, redirect chains and crawler access, and xml sitemap audit checklist.
Frequently asked questions
Do AI crawlers need separate redirect rules from search engine crawlers?
Not necessarily separate rules, but the same redirect must actually work for non-browser, non-JavaScript-executing clients, which some migrations overlook.
How long should old-URL monitoring continue after migration?
Monitor until AI-crawler-identified traffic to old paths drops to a low, stable baseline; the exact duration varies by site and crawler re-discovery rate, so treat any fixed number as a rough estimate rather than a guarantee.