Agent ReadyAI shopping agent readiness

How to make your store AI-crawler friendly

Part of the Agent Ready guides. Plain advice, no signup needed.

An AI crawler is a program that reads web pages on behalf of an AI assistant. When a shopper asks an assistant where to buy a product, the assistant works from what these crawlers have been able to read. If a crawler cannot read your store, the assistant has nothing to quote and moves on to a store it can read.

The main crawlers publish their names. GPTBot reads pages used in ChatGPT answers. ChatGPT-User fetches pages live when a user asks about a specific link. ClaudeBot reads pages used in Claude answers. PerplexityBot reads pages used in Perplexity answers. Applebot-Extended, Bytespider, and CCBot collect training and index data that can also feed answers. The names change over time, so treat any list as a starting point and check your own server logs for who actually visits.

Being crawler friendly means three things. Your server answers crawler requests with the same content a browser gets. Your robots file permits the crawlers you want. Your pages carry the facts assistants quote: name, price, availability, shipping, and returns. This guide walks through each one.

Step 1: See who can reach your store today

Start with evidence, not guesses. Open your server or CDN logs for the last 30 days and search for the crawler names above. Note three things for each: how many times it visited, which status codes it got, and which pages it asked for. A crawler that never appears may be blocked. A crawler that gets mostly 403 or 429 responses is being blocked. A crawler that only ever sees the homepage cannot learn your catalog.

If you have no log access, use what your host gives you. Most managed hosts keep a traffic or access log in the control panel. Cloudflare users can check the firewall events log for blocked requests by user agent. Shopify merchants do not get raw logs, so skip to the robots and page checks below and rely on the free audit instead.

Fetch your own robots file at yourstore.com/robots.txt and read every line. Look for any Disallow rule naming an AI crawler, and any blanket Disallow rule. Then fetch a product page with a plain text tool such as curl and compare it with what a browser shows. If the price or the add to cart button is missing from the plain fetch, part of your page renders only in JavaScript, which some crawlers handle poorly. That gap is fixed in step 3.

Step 2: Permit the crawlers in robots.txt

robots.txt sits at your domain root and tells crawlers which paths they may visit. AI crawlers obey it. One wrong line can hide your whole store from one assistant while everything looks fine in a browser. Keep the file short and explicit.

User-agent: GPTBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: anthropic-ai
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Applebot-Extended
Allow: /

User-agent: CCBot
Allow: /

User-agent: *
Allow: /

Place specific crawler blocks before the wildcard block. A crawler uses the block that names it and ignores the rest, so a permissive wildcard does not rescue a crawler that has its own restrictive block elsewhere in the file. After editing, fetch the file again in a private browser window to confirm the live version changed. Cached copies can linger for a day on some hosts.

Some merchants want search indexing but not AI training data. Published crawler docs describe separate agents for separate jobs: GPTBot collects training data while ChatGPT-User fetches live pages for answers, and Google-Extended controls AI features separately from Googlebot search crawling. If that distinction matters to you, restrict the training collector and permit the live fetching agent, then watch your logs to confirm each behaves as expected. When in doubt, permit both. A blocked training crawler costs you nothing visible today, but a blocked answering crawler costs you recommendations.

Step 3: Serve the same content to crawlers and browsers

The most common silent failure is a page that looks complete in a browser but arrives nearly empty to a crawler. Product grids that load on scroll, prices injected by JavaScript after the page loads, and reviews pulled in by a third party script are the usual causes. The fix is server rendering for the facts that matter: product name, price, currency, availability, main image, and a short description should be present in the first HTML response.

Test this yourself. Turn off JavaScript in your browser and reload a product page. Everything an assistant needs should still be visible: the name, the price, the stock state, and the shipping or returns link. If the page falls apart with scripts off, treat that as the punch list. You do not need to remove interactive features. You need the core facts in the HTML.

Keep pages fast for crawlers too. Crawlers work through a visit budget per site. A product page that takes six seconds to answer eats budget that could have covered ten more pages. Compress images, keep third party scripts to the ones that earn their place, and make sure the server answers with caching headers so repeat visits are cheap. Aim for the main content to arrive in the first response, not after a chain of redirects. Each redirect is another round trip the crawler may not wait for.

Step 4: Point crawlers at your catalog

Discovery is the crawler finding your products in the first place. Keep an XML sitemap that lists every product page and refreshes when products are added or removed. Reference it from robots.txt with a Sitemap line so a crawler that has never seen your store can start there. Check that the sitemap returns a 200 status and valid XML; a sitemap that errors out is worse than none because it wastes the visit.

Keep category pages linked and crawlable. A category page that filters products only through JavaScript, with no plain links to the products, is a dead end for a simple crawler. Every product should be reachable by plain links within a few clicks of the homepage. Paginated categories should use plain links for page two, page three, and so on, not buttons that only work with scripts.

If you sell through Google surfaces, keep a product feed in Merchant Center with prices and availability that match your pages. Assistants cross check feed data against page data. A feed that says in stock while the page says out of stock teaches every consumer of that data not to trust you. Match the two exactly, including currency and any price suffixes such as VAT included.

Step 5: Remove the blocks you did not know you had

Bot managers and firewall rules are the next place to look. Many security plugins and CDNs ship with bot blocking presets that predate AI crawlers and lump them in with scrapers. Open your firewall and rate limiting settings and search for the crawler names. If any are set to block or challenge, change them to allow and monitor for a week. A challenge page, such as a CAPTCHA or a JavaScript proof of work page, reads as a wall to a crawler. Humans never see it, so it can sit there for months while assistants see nothing but the wall.

Rate limits deserve care. Aggressive throttling that returns 429 responses to anything fast looks identical to blocking from the crawler side. Set limits that stop abuse without punishing legitimate crawlers: per address limits in the range your host recommends, with known crawler agents exempted or given generous quotas. Then confirm in the logs that crawler visits complete with 200 responses.

Check IP level blocks too. If your store or your host blocks whole countries or address ranges, confirm none of your important crawlers fetch from those ranges. When a crawler fails from one region, some systems retry from another, but not all do. Country blocking is a business decision; just make it knowingly, with the crawling cost counted.

Common mistakes

Copying a robots template that blocks GPTBot or CCBot by default. This is the single most frequent cause of an invisible store, and it usually arrives pasted inside an SEO starter template the merchant never read. Read your own file line by line.

Blocking staging rules on the live site. A Disallow for the whole site makes sense on a staging copy and is catastrophic on production. After every migration or redesign, fetch the live robots file and confirm it is the permissive one.

Hiding prices behind logins or region pickers. A crawler does not log in and does not pick a region. If the price only appears after both, assistants will quote your products without prices or skip them. Show list prices to anonymous visitors.

Serving different content per user agent. Showing crawlers a stripped page while browsers get the full page is cloaking. Search engines penalize it, and assistants that compare sources will distrust the mismatch. Serve everyone the same facts.

Forgetting the mobile page. Some crawlers fetch your mobile rendering. If your mobile theme drops product descriptions or structured data that the desktop theme has, fix the mobile theme rather than maintaining two truths.

On Shopify

Shopify controls the server layer, so steps about logs and firewalls mostly do not apply. Your levers are the theme, robots template, sitemap, and apps. Edit robots.txt.liquid under Online Store, Themes, the theme menu, Edit code, Templates. Keep your additions minimal and recheck the live file after every theme update, since theme changes can overwrite the template.

Shopify generates your sitemap automatically at /sitemap.xml, so discovery is handled as long as products are published to the online store channel. Unpublished or draft products never appear in it. Check that products you want recommended are published, in stock, and assigned to the right collections with plain links.

Theme choice matters for step 3. Older or heavily customized themes sometimes render prices and variants only through scripts. Test with scripts off as described above. If facts vanish, ask the theme developer for server rendered product facts, or switch the product template to one that prints them in HTML. Schema apps can add markup but cannot fix content that never reaches the HTML.

On WooCommerce

You control the server, so work through all five steps in order. Start with the SEO plugin file editor: Rank Math and Yoast both expose robots.txt editing under their settings, and either will warn about syntax errors. Keep a backup of the file before each change so a bad edit takes seconds to undo.

Cache plugins are the usual suspect for empty crawler pages. Some cache setups serve a minimal cached shell and fill prices in per visitor with scripts. Exclude product price and stock fragments from that behavior, or switch the cache to full page caching that still contains the facts. Then verify with curl that a product URL returns name, price, and availability in the HTML.

Security plugins need an allowlist pass. Wordfence, Sucuri, and similar tools can challenge unrecognized agents. Add the AI crawler agents to the allowlist and set rate limits that humans will never notice but crawlers can work within. Recheck after every plugin update, since major updates sometimes reset custom rules.

How to confirm it worked

Verification is a loop, not a single check. After each change, fetch robots.txt and one product page with curl using a crawler user agent string and confirm a 200 response with the facts present. Then watch the logs for two weeks: crawler visits should rise, 403 and 429 responses should fall toward zero, and the pages visited should spread across your catalog instead of clustering on the homepage.

Run an outside audit as a second pair of eyes. Automated checks catch the file level problems, missing fields, and blocked agents in minutes, while logs tell you the longer story. Fix what the audit flags first, since those are the prerequisites, then use the logs to confirm real crawlers now get through.

Want to know where your store stands today. Run the free Agent Ready audit to get a graded score and the exact fixes that matter most for your store. It takes under a minute and needs no signup.

Run free audit

Get a free AI-crawler audit of your store

Enter your email and store address and we will keep you posted on your score. After signup you can run the audit right away.

More guides: ChatGPT Shopping Shopify schema robots.txt for AI AI-crawler friendly Product schema llms.txt AI bot rules FAQ schema