Search crawlers and AI crawlers are not the same thing
Googlebot and Bingbot build the search indexes that send you shoppers. They are not the problem this guide is about, and they must stay allowed. If you block them, they stop crawling your store and your pages fall out of normal search results. Google notes that a page it is not allowed to crawl can still show up, but only as a link without a description.
Never block Googlebot or Bingbot to deal with AI crawlers, in robots.txt or on the server. Google's AI Overviews and AI Mode are part of Google Search and use Googlebot, so blocking Googlebot to avoid them removes your store from Google Search as well.
AI crawlers do three different jobs, and the difference matters when you decide what to block:
- Training: collecting content that may be used to train AI models, such as OpenAI's GPTBot, Anthropic's ClaudeBot and Meta's meta-externalagent. Google and Apple use robots.txt tokens rather than separate crawlers for this: Google-Extended and Applebot-Extended.
- AI search and answers: finding pages to show or cite in AI search, such as OAI-SearchBot, Claude-SearchBot, PerplexityBot and meta-webindexer. Block these and your store can appear less in those answers.
- User-initiated fetches: visiting a page because a user asked the assistant to, such as ChatGPT-User, Claude-User, Perplexity-User and meta-externalfetcher. Several of their owners say robots.txt may not apply to these.
The AI crawlers worth knowing
Every name below comes from the owner's own documentation. The last column is what the owner says about robots.txt. We have left out crawlers whose owners do not publish clear documentation.
| User agent | Owner | What it does | Follows robots.txt? |
|---|---|---|---|
| GPTBot | OpenAI | Collects content that may be used to train OpenAI's models | Yes |
| OAI-SearchBot | OpenAI | Finds pages to show in ChatGPT search. Not used for training | Yes |
| ChatGPT-User | OpenAI | Visits pages when a ChatGPT user asks it to | May not apply |
| ClaudeBot | Anthropic | Collects content that could be used to train Claude models | Yes |
| Claude-SearchBot | Anthropic | Crawls to improve Claude's search results | Yes |
| Claude-User | Anthropic | Fetches pages when a Claude user asks | Yes |
| Google-Extended | Not a crawler: a robots.txt token for Gemini training and grounding. No effect on Google Search | Token only | |
| Applebot-Extended | Apple | Not a crawler: opts content out of training Apple's foundation models. Pages can still appear in Apple's search results | Token only |
| PerplexityBot | Perplexity | Surfaces and links sites in Perplexity search. Not used for training | Yes |
| Perplexity-User | Perplexity | Fetches pages to answer a user's question | Generally ignores it |
| meta-externalagent | Meta | Crawls for uses such as training foundation AI models | Yes |
| meta-webindexer | Meta | Improves Meta AI search results | Yes |
| meta-externalfetcher | Meta | Fetches individual links at a user's request | May bypass it |
| CCBot | Common Crawl | Builds Common Crawl's open web archive, which anyone can use | Blockable by name |
Decide what you actually want
There is no single right answer, so it is worth being clear about the trade-off before you change anything.
- Block training only. Your content stays out of future model training, while AI search and answer tools can still find and cite your products. Google Search and Bing are unaffected.
- Block every AI crawler. The strongest opt-out, but your store can appear less in AI answers, which some shoppers now use to find products. OpenAI, for example, says sites that block OAI-SearchBot will not be shown in ChatGPT search answers.
- Slow them down instead. If the real issue is load rather than principle, rate limiting at the server keeps the bots but caps how hard they can hit you. Anthropic also supports the non-standard Crawl-delay line for its crawlers; Google's crawlers ignore it.
Example robots.txt rules
In the Magento admin, go to Content > Design > Configuration, choose the scope, open Search Engine Robots and use the Edit custom instruction of robots.txt file field. Keep Adobe's default instructions and add your AI crawler groups underneath. Then open /robots.txt on your store and check what it actually serves: if a robots.txt file already sits on the server, that may be the one visitors get.
Option 1, block training only:
# Block AI training only. Search and AI answer crawlers are still allowed.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: meta-externalagent
Disallow: /
# Optional: Common Crawl's public archive
User-agent: CCBot
Disallow: /Option 2, block every AI crawler: use the rules above plus these.
# Add these to the training-only rules to block AI search and answer bots too.
User-agent: OAI-SearchBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
User-agent: Claude-SearchBot
Disallow: /
User-agent: Claude-User
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: Perplexity-User
Disallow: /
User-agent: meta-webindexer
Disallow: /
User-agent: meta-externalfetcher
Disallow: /One rule of robots.txt catches people out. Google and Bing both say a crawler follows only the most specific group that names it and ignores the rest, so a group written for one bot does not inherit your User-agent: * rules. That is harmless for the Disallow: / groups above. It matters if you ever write a group for Googlebot or Bingbot, because you would then need to repeat your default rules inside it.
robots.txt is a request, not a lock
Google's own documentation is plain about this: robots.txt cannot enforce crawler behaviour, and it is up to each crawler to obey it. The major AI companies say their crawlers do, but the user-initiated fetchers above may not, and some bots fake their user agent. Common Crawl, for example, warns about crawlers falsely identifying themselves as CCBot.
The backstop is the web server. Your host or developer can refuse requests from named user agents, or rate limit them, before they reach Magento. On nginx that can be as simple as this, as general advice to adapt rather than paste:
# Example only: refuse named AI crawlers at the web server.
# Never add Googlebot or bingbot to this list.
map $http_user_agent $block_ai_bot {
default 0;
~*(GPTBot|ClaudeBot|meta-externalagent|CCBot) 1;
}
server {
# ...your existing Magento server block...
if ($block_ai_bot) { return 403; }
}Blocking by user agent only works on bots that tell the truth about who they are. For the rest, blocking or rate limiting by IP address is the next step, and several owners publish their IP ranges so you can check a claim before acting on it: OpenAI, Anthropic and Perplexity publish lists of the addresses their crawlers use, Common Crawl runs CCBot from dedicated ranges with reverse DNS, and Google documents how to confirm that a visitor claiming to be Googlebot really is. Check before you block anything that says it is Googlebot.
Why crawlers can weigh on a Magento store
This part is general advice rather than anything specific to one bot. A Magento storefront with layered navigation can produce a very large number of filter URL combinations, and a crawler that follows them all keeps asking for pages that have never been requested before. Varnish serves repeat requests for the same page from memory, but each new combination has to be built by PHP and the database the first time, and on-site search results work the same way. That is why a busy crawler can cost a Magento store more than its share of visits suggests.
Canonical tags and a considered robots.txt help, but be careful about disallowing filter URLs for every crawler: those rules apply to Google too, so decide what you want indexed first. If you are sizing a server, measure bot traffic from the server side, because analytics tools usually do not count it.
Where your hosting helps
Which AI crawlers may read your store is your decision, and robots.txt is where you make it. Hosting helps with the bots that do not listen. On EveryHost plans, Sentinel security monitoring reads the web server logs and can automatically block abusive IP addresses at the server firewall, and our engineers review every block. The same engineers can help you add user agent or rate limiting rules on the server when a crawler ignores robots.txt, without touching Googlebot or Bingbot.
Aggressive crawlers are not the only automated traffic worth watching. Bots trying stolen card numbers at your checkout are a different problem with different fixes, covered in our guide to stopping card testing on Magento 2. If you would rather have a UK team who work only on Magento looking after all of this, our managed Magento VPS hosting includes Sentinel on every plan.
Frequently Asked Questions
Sources and further reading
- OpenAI, Overview of OpenAI crawlers
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Google, Google's common crawlers (Google-Extended) and AI features and your website
- Google, Introduction to robots.txt, How Google interprets the robots.txt specification, and Verifying Google crawlers
- Apple, About Applebot
- Perplexity, Perplexity crawlers
- Meta, Meta web crawlers
- Common Crawl, CCBot
- Bing Webmaster Tools, Which crawlers does Bing use and How to create a robots.txt file
- Adobe Experience League, Search engine robots
Crawlers slowing your store down?
Our UK Magento engineers can look at your server logs with you, set up robots.txt the way you want it and add server rules for bots that ignore it, without touching Google or Bing. Sentinel security monitoring is included on every plan.