Agent ReadyAI shopping agent readiness

Robots.txt for AI bots

Part of the Agent Ready guides. Plain advice, no signup needed.

robots.txt is a plain text file at your domain root that tells automated visitors which parts of your site they may fetch. Every serious crawler reads it first, including the AI bots that feed shopping assistants. A correct file takes ten minutes to write and protects every product page you own. A wrong one hides your store from an entire assistant while your search traffic looks perfectly normal.

This guide covers the AI bots by name, the exact rules for each situation, and how to test the result. It assumes you already know the file exists. If you want the short version first, read our robots.txt guide for AI crawlers, then come back here for the detailed rules.

How the file is read

The file is a list of blocks. Each block starts with one or more User-agent lines naming a crawler, followed by Allow and Disallow rules for that crawler. A crawler finds the block that names it and follows only that block. Crawlers with no named block follow the wildcard block, the one labeled User-agent with a star. This matching order is the source of nearly every robots mistake: a specific block always beats the wildcard, even when the wildcard is more permissive.

Rules are prefix matches against the path. Disallow: /blocks everything under that path. Allow: / permits everything. An empty Disallow with nothing after the colon permits everything in that block. Comments start with a hash sign and are ignored. The file is case sensitive in paths, so /Sale and /sale are different paths to a crawler.

One file serves the whole host. The file at yourstore.com/robots.txt governs only that exact host: protocol, domain, and port all matter. A store with separate hosts for regions or languages needs a correct file on each host. Subdomains do not inherit the main domain file.

The AI bot names worth knowing

Crawler operators publish the user agent strings their bots send. The shopping relevant ones are GPTBot, which gathers browsing data for ChatGPT features, and ChatGPT-User, which fetches specific pages live when a user pastes a link or asks about one. ClaudeBot and anthropic-ai serve Claude features. PerplexityBot serves Perplexity answers. Applebot-Extended feeds Apple intelligence features. Bytespider collects data used in TikTok parent company models. CCBot maintains the Common Crawl dataset that many models train on. Google-Extended governs AI feature use of your pages separately from Googlebot search crawling. Meta-WebIndexer and Meta-ExternalAgent serve Meta AI features.

You do not need a block for every name on that list. The practical approach is to name the bots you have a specific decision about and let the wildcard cover the rest. Most stores end up naming five to eight agents explicitly: the shopping answer bots they want to permit, plus any training collectors they want to restrict.

The permissive file: allow the answer bots

Start from open. The default below permits every named AI bot plus everything else, while keeping admin and checkout internals out of crawling. Adjust the internal paths to match your platform: cart, checkout, and account paths differ between Shopify, WooCommerce, and custom builds.

User-agent: GPTBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: anthropic-ai
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Applebot-Extended
Allow: /

User-agent: CCBot
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: *
Allow: /
Disallow: /admin/
Disallow: /cart
Disallow: /checkout
Disallow: /account

Keep the internal Disallow rules in the wildcard block only, unless a named bot also needs them. Repeating them per bot adds lines without adding protection, since named bots that you permit should still skip checkout flows. Note that Disallow is a crawling instruction, not access control: it asks polite bots not to fetch, but it does not password protect anything. Private pages need real authentication regardless of what robots says.

The selective file: restrict training, permit answers

Some merchants are comfortable with assistants quoting their pages but do not want their catalog used as model training data. Published operator docs describe separate agents for these jobs, which makes a split file possible. The pattern below restricts the collectors while permitting the live answer fetchers. Verify the agent roles against current operator docs before deploying this, because operators rename agents and the roles are theirs to define.

User-agent: GPTBot
Disallow: /

User-agent: ChatGPT-User
Allow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: PerplexityBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: *
Allow: /

Understand the tradeoff before you use it. Restricting a training collector can reduce how often your products surface in answers that draw on training data, while permitting the fetcher keeps live lookups working. There is no public meter that shows the split, so watch your own referral and mention data after the change. When the goal is maximum visibility to shoppers, the permissive file is the safer choice.

Path rules for store internals

Keep crawlers out of paths that waste their visit budget or create junk URLs. Cart, checkout, account, wishlist, and internal search result pages all qualify. A crawler that spends its visit on a thousand search result variations has less budget left for product pages. Disallow the search results path and the filtered faceted URLs if your platform generates them, while keeping category and product paths open.

Pagination needs care. Allowing page two, page three, and deeper category pages is correct when those pages link to products. Blocking pagination to save budget also blocks the products that only appear on later pages. If your catalog is large, prefer a complete sitemap plus open pagination over blocking either.

Staging and development hosts need the opposite file: everything disallowed. A staging copy that search engines or assistants can reach creates duplicate content and leaks unreleased products. The staging file is two lines: the wildcard agent and Disallow for the root. Confirm it on every staging deploy, and confirm the production file is the permissive one after every launch.

Common mistakes

A named block that contradicts the wildcard. A template that blocks one AI bot by name while the wildcard allows everything leaves that bot blocked. The specific block wins. Read every named block on its own terms instead of assuming the wildcard rescues it.

Trailing slash confusion. Disallow: /cart blocks /cart and everything under it but not /cartoon. Allow and Disallow are prefix matches, so check that short rules do not accidentally cover longer paths you want open, and that rules meant to cover a section include the trailing slash.

Case errors in paths. /Checkout and /checkout are different rules. Match the exact casing your platform uses. Shopify paths are lowercase; custom builds may not be. Copy paths from the live URL, not from memory.

Editing the wrong host file. Stores with www and apex variants, or with regional subdomains, need the fix applied per host. Check the file on each hostname your store answers on. Redirects between hosts do not carry robots rules along.

Syntax the parser cannot use. A missing colon after User-agent, a rule with no path and no Allow intent, or non ASCII characters pasted from a rich text editor can void a line or a block. Write the file in a plain text editor and validate it after every change.

Assuming the file updates instantly. Crawlers cache the file, some for up to a day. After fixing a block, expect stragglers that still obey the old rules for a while. Confirm the live file is correct, then give it 24 to 48 hours before judging crawler behavior.

How to test before it costs you

Test in three layers. First, fetch the live file and read it line by line, checking each named block against your intent. Second, run it through a robots testing tool, which shows exactly which paths each named agent may fetch. Third, fetch a product page with curl using each important crawler user agent string and confirm a 200 response with the product facts present.

Re-test on a schedule, not just after edits. Theme updates, migrations, SEO plugin changes, and staging syncs all overwrite or alter the file. A monthly fetch of the live file takes a minute and catches the regression before it costs a month of crawler visits. Add it to the same checklist as your backup test.

On Shopify

Edit robots.txt.liquid under Online Store, Themes, the theme menu, Edit code, Templates. Shopify wraps your rules with its own generated content, so preview the live file after saving and confirm your blocks appear verbatim. Theme updates can replace the template, so recheck the live file after every theme change.

Shopify path defaults to disallow are handled by the platform template: admin, cart, and checkout paths are already restricted in the default output. Your job is adding the AI bot blocks and confirming nothing in the generated portion blocks them. Keep your additions in one clearly marked section so the next person editing the file sees the intent.

On WooCommerce

Edit through your SEO plugin: Rank Math hides it under General Settings, Files Editor, and Yoast under Tools, File Editor. Both validate syntax on save. If file editing is disabled by your host, often via a DISALLOW_FILE_EDIT constant, use SFTP and edit the physical robots.txt at the site root instead.

Watch for plugin conflicts. Caching plugins, security plugins, and multilingual plugins can each modify robots output. After saving, fetch the live file in a private window and confirm it contains exactly your rules plus what you expect from each plugin. If rules appear that you did not write, find the plugin that adds them before deleting anything.

WooCommerce internal paths to keep disallowed include cart, checkout, my-account, and add to cart query URLs. Product category and tag archives stay open. If you use faceted filter plugins that generate indexed URLs, disallow the filter parameter patterns while keeping the base category open.

Want the rest of your store checked the same way. Run the free Agent Ready audit to get a graded score and the exact fixes that matter most for your store. It takes under a minute and needs no signup.

Run free audit

Get a free AI-crawler audit of your store

Enter your email and store address and we will keep you posted on your score. After signup you can run the audit right away.

More guides: ChatGPT Shopping Shopify schema robots.txt for AI AI-crawler friendly Product schema llms.txt AI bot rules FAQ schema