How-to guide

    How to Block AI Crawlers on Magento 2

    Without Hurting Your Google or Bing Rankings

    By Simon Bumford, Founder of EveryHost••10 min read

    TL;DR

    AI companies send several different crawlers: some collect content for training, some power AI search and answers, and some fetch a page when a user asks. You can control most of them with a few lines in robots.txt, which Magento lets you edit from the admin. Never block Googlebot or Bingbot to do it, because that takes your store out of normal search results. Decide whether you want to block training only or every AI crawler, knowing that blocking AI search bots can mean appearing less in AI answers. robots.txt is a request rather than a lock, so the backstop for bots that ignore it is blocking or rate limiting at the server.

    Search crawlers and AI crawlers are not the same thing

    Googlebot and Bingbot build the search indexes that send you shoppers. They are not the problem this guide is about, and they must stay allowed. If you block them, they stop crawling your store and your pages fall out of normal search results. Google notes that a page it is not allowed to crawl can still show up, but only as a link without a description.

    Never block Googlebot or Bingbot to deal with AI crawlers, in robots.txt or on the server. Google's AI Overviews and AI Mode are part of Google Search and use Googlebot, so blocking Googlebot to avoid them removes your store from Google Search as well.

    AI crawlers do three different jobs, and the difference matters when you decide what to block:

    • Training: collecting content that may be used to train AI models, such as OpenAI's GPTBot, Anthropic's ClaudeBot and Meta's meta-externalagent. Google and Apple use robots.txt tokens rather than separate crawlers for this: Google-Extended and Applebot-Extended.
    • AI search and answers: finding pages to show or cite in AI search, such as OAI-SearchBot, Claude-SearchBot, PerplexityBot and meta-webindexer. Block these and your store can appear less in those answers.
    • User-initiated fetches: visiting a page because a user asked the assistant to, such as ChatGPT-User, Claude-User, Perplexity-User and meta-externalfetcher. Several of their owners say robots.txt may not apply to these.

    The AI crawlers worth knowing

    Every name below comes from the owner's own documentation. The last column is what the owner says about robots.txt. We have left out crawlers whose owners do not publish clear documentation.

    User agentOwnerWhat it doesFollows robots.txt?
    GPTBotOpenAICollects content that may be used to train OpenAI's modelsYes
    OAI-SearchBotOpenAIFinds pages to show in ChatGPT search. Not used for trainingYes
    ChatGPT-UserOpenAIVisits pages when a ChatGPT user asks it toMay not apply
    ClaudeBotAnthropicCollects content that could be used to train Claude modelsYes
    Claude-SearchBotAnthropicCrawls to improve Claude's search resultsYes
    Claude-UserAnthropicFetches pages when a Claude user asksYes
    Google-ExtendedGoogleNot a crawler: a robots.txt token for Gemini training and grounding. No effect on Google SearchToken only
    Applebot-ExtendedAppleNot a crawler: opts content out of training Apple's foundation models. Pages can still appear in Apple's search resultsToken only
    PerplexityBotPerplexitySurfaces and links sites in Perplexity search. Not used for trainingYes
    Perplexity-UserPerplexityFetches pages to answer a user's questionGenerally ignores it
    meta-externalagentMetaCrawls for uses such as training foundation AI modelsYes
    meta-webindexerMetaImproves Meta AI search resultsYes
    meta-externalfetcherMetaFetches individual links at a user's requestMay bypass it
    CCBotCommon CrawlBuilds Common Crawl's open web archive, which anyone can useBlockable by name

    Decide what you actually want

    There is no single right answer, so it is worth being clear about the trade-off before you change anything.

    • Block training only. Your content stays out of future model training, while AI search and answer tools can still find and cite your products. Google Search and Bing are unaffected.
    • Block every AI crawler. The strongest opt-out, but your store can appear less in AI answers, which some shoppers now use to find products. OpenAI, for example, says sites that block OAI-SearchBot will not be shown in ChatGPT search answers.
    • Slow them down instead. If the real issue is load rather than principle, rate limiting at the server keeps the bots but caps how hard they can hit you. Anthropic also supports the non-standard Crawl-delay line for its crawlers; Google's crawlers ignore it.

    Example robots.txt rules

    In the Magento admin, go to Content > Design > Configuration, choose the scope, open Search Engine Robots and use the Edit custom instruction of robots.txt file field. Keep Adobe's default instructions and add your AI crawler groups underneath. Then open /robots.txt on your store and check what it actually serves: if a robots.txt file already sits on the server, that may be the one visitors get.

    Option 1, block training only:

    # Block AI training only. Search and AI answer crawlers are still allowed.
    User-agent: GPTBot
    Disallow: /
    
    User-agent: ClaudeBot
    Disallow: /
    
    User-agent: Google-Extended
    Disallow: /
    
    User-agent: Applebot-Extended
    Disallow: /
    
    User-agent: meta-externalagent
    Disallow: /
    
    # Optional: Common Crawl's public archive
    User-agent: CCBot
    Disallow: /

    Option 2, block every AI crawler: use the rules above plus these.

    # Add these to the training-only rules to block AI search and answer bots too.
    User-agent: OAI-SearchBot
    Disallow: /
    
    User-agent: ChatGPT-User
    Disallow: /
    
    User-agent: Claude-SearchBot
    Disallow: /
    
    User-agent: Claude-User
    Disallow: /
    
    User-agent: PerplexityBot
    Disallow: /
    
    User-agent: Perplexity-User
    Disallow: /
    
    User-agent: meta-webindexer
    Disallow: /
    
    User-agent: meta-externalfetcher
    Disallow: /

    One rule of robots.txt catches people out. Google and Bing both say a crawler follows only the most specific group that names it and ignores the rest, so a group written for one bot does not inherit your User-agent: * rules. That is harmless for the Disallow: / groups above. It matters if you ever write a group for Googlebot or Bingbot, because you would then need to repeat your default rules inside it.

    robots.txt is a request, not a lock

    Google's own documentation is plain about this: robots.txt cannot enforce crawler behaviour, and it is up to each crawler to obey it. The major AI companies say their crawlers do, but the user-initiated fetchers above may not, and some bots fake their user agent. Common Crawl, for example, warns about crawlers falsely identifying themselves as CCBot.

    The backstop is the web server. Your host or developer can refuse requests from named user agents, or rate limit them, before they reach Magento. On nginx that can be as simple as this, as general advice to adapt rather than paste:

    # Example only: refuse named AI crawlers at the web server.
    # Never add Googlebot or bingbot to this list.
    map $http_user_agent $block_ai_bot {
        default 0;
        ~*(GPTBot|ClaudeBot|meta-externalagent|CCBot) 1;
    }
    
    server {
        # ...your existing Magento server block...
        if ($block_ai_bot) { return 403; }
    }

    Blocking by user agent only works on bots that tell the truth about who they are. For the rest, blocking or rate limiting by IP address is the next step, and several owners publish their IP ranges so you can check a claim before acting on it: OpenAI, Anthropic and Perplexity publish lists of the addresses their crawlers use, Common Crawl runs CCBot from dedicated ranges with reverse DNS, and Google documents how to confirm that a visitor claiming to be Googlebot really is. Check before you block anything that says it is Googlebot.

    Why crawlers can weigh on a Magento store

    This part is general advice rather than anything specific to one bot. A Magento storefront with layered navigation can produce a very large number of filter URL combinations, and a crawler that follows them all keeps asking for pages that have never been requested before. Varnish serves repeat requests for the same page from memory, but each new combination has to be built by PHP and the database the first time, and on-site search results work the same way. That is why a busy crawler can cost a Magento store more than its share of visits suggests.

    Canonical tags and a considered robots.txt help, but be careful about disallowing filter URLs for every crawler: those rules apply to Google too, so decide what you want indexed first. If you are sizing a server, measure bot traffic from the server side, because analytics tools usually do not count it.

    Where your hosting helps

    Which AI crawlers may read your store is your decision, and robots.txt is where you make it. Hosting helps with the bots that do not listen. On EveryHost plans, Sentinel security monitoring reads the web server logs and can automatically block abusive IP addresses at the server firewall, and our engineers review every block. The same engineers can help you add user agent or rate limiting rules on the server when a crawler ignores robots.txt, without touching Googlebot or Bingbot.

    Aggressive crawlers are not the only automated traffic worth watching. Bots trying stolen card numbers at your checkout are a different problem with different fixes, covered in our guide to stopping card testing on Magento 2. If you would rather have a UK team who work only on Magento looking after all of this, our managed Magento VPS hosting includes Sentinel on every plan.

    Frequently Asked Questions

    Not if you block the right ones. Blocking AI training crawlers such as GPTBot, ClaudeBot or the Google-Extended token does not affect Google Search: Google says Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal. What does hurt is blocking Googlebot or Bingbot, which stops them crawling your store and takes your pages out of normal search results.

    No. Google says AI Overviews and AI Mode are part of Search, so they are controlled through Googlebot and the usual preview controls such as nosnippet and max-snippet. Google-Extended controls whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and Vertex AI. It does not change how your store appears in Google Search.

    GPTBot collects content that may be used to train OpenAI's models. OAI-SearchBot finds pages to show in ChatGPT's search features and is not used for training. OpenAI treats each setting separately, so you can block GPTBot and still allow OAI-SearchBot. OpenAI says sites that block OAI-SearchBot will not be shown in ChatGPT search answers.

    No. robots.txt is a request, and Google itself says the file cannot enforce crawler behaviour. The major AI companies say their crawlers follow it, but fetchers that act on a user's request, such as ChatGPT-User, Perplexity-User and Meta-ExternalFetcher, may not, and some bots fake their user agent. For those, the backstop is blocking or rate limiting at the server.

    In the Magento admin, go to Content > Design > Configuration, choose the scope, open Search Engine Robots and add your rules in the Edit custom instruction of robots.txt file field. Keep Adobe's default instructions, add your AI crawler groups below them, then open /robots.txt on your store to check the result. If a robots.txt file already sits on the server, ask your developer or host which one your store is actually serving.

    That depends on what you want. Blocking training crawlers keeps your content out of future model training without affecting search. Blocking AI search and answer crawlers as well can mean your store appears less in AI answers, which some shoppers now use to find products. If crawlers are straining your server rather than worrying you on principle, rate limiting at the server can be a better fit than an outright block.

    Sources and further reading

    • OpenAI, Overview of OpenAI crawlers
    • Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?
    • Google, Google's common crawlers (Google-Extended) and AI features and your website
    • Google, Introduction to robots.txt, How Google interprets the robots.txt specification, and Verifying Google crawlers
    • Apple, About Applebot
    • Perplexity, Perplexity crawlers
    • Meta, Meta web crawlers
    • Common Crawl, CCBot
    • Bing Webmaster Tools, Which crawlers does Bing use and How to create a robots.txt file
    • Adobe Experience League, Search engine robots

    Crawlers slowing your store down?

    Our UK Magento engineers can look at your server logs with you, set up robots.txt the way you want it and add server rules for bots that ignore it, without touching Google or Bing. Sentinel security monitoring is included on every plan.