Is Cloudflare blocking AI shoppers from your store?
Part of the Agent Ready guides. Plain advice, no signup needed.
If your store sits behind Cloudflare, a change made on September 15, 2026 may be blocking AI shopping assistants from reading your pages. Cloudflare changed its defaults so that two AI crawler groups, called Training and Agent, are now blocked by default for new domains, for sites that existing customers add from now on, and for free tier accounts. Search stays allowed. This guide explains what changed, why it matters for stores, how to check your site, and how to fix the setting.
The short version is this. The Agent group includes the fetchers that retrieve pages for AI assistants and agents, which includes tools that act for a buyer researching a purchase. If your domain falls under the new defaults, those fetchers can be stopped at the Cloudflare edge before they reach your origin server. Your own robots file can look clean while the edge still refuses the request.
What changed on September 15
On September 15, 2026, Cloudflare started to block the Training and Agent crawler categories by default for new domains, for new sites added by existing customers, and for free tier accounts. Older paid setups keep their prior settings unless the owner changes them. Search stays allowed by default, so normal search indexing is not affected.
Cloudflare now treats Search, Training, and Agent as separate controls. Search covers crawlers that index pages for search results. Training covers crawlers that collect text for model training. Agent covers crawlers that retrieve pages for AI assistants and agents at answer time. Before, many owners used one broad control for AI bots. Now each category has its own switch, and Training and Agent default to block in the affected accounts.
The block is enforced at the network layer, not through robots.txt. Robots.txt is a text file that asks polite crawlers to stay away from listed paths. A network layer block refuses the request at the edge, before origin rules are read. Even if your robots file permits a crawler by name, the edge block can still stop it. Merchants who check only their own file miss this.
Cloudflare can also inject its own managed robots block into the robots file that crawlers fetch. When the managed robots option is on, the served file holds a Cloudflare written section above your own rules. That section can list AI crawlers with Disallow rules that sit above the Allow rules you wrote below. The file you maintain at the origin is still there, but the served file is the combination, and the managed part comes first.
Why it matters for a store
Many Shopify and direct to consumer stores sit behind Cloudflare, often without the merchant thinking about it. The domain may have been added for speed or security, or by an agency during setup, then left alone. If that domain was added recently, or if the account is on the free tier, the new defaults may apply. The owner changed nothing, yet AI shopping agents started to see refusals.
A store in this state silently blocks AI shopping agents. Pages load for shoppers. Search shows indexed pages. Ads run. Nothing in daily work looks wrong. But when an assistant tries to fetch a product page to answer a buyer question, the edge refuses the fetch. The assistant then quotes a store it could read. The loss never appears as a visit, because the blocked fetch never becomes one.
Edits at the origin cannot remove an edge injected block. If you edit robots.txt.liquid in Shopify, or the physical file on WooCommerce, you change only your own rules. You do not change the managed section above them, and you do not change the network layer setting. Both stay until someone with Cloudflare dashboard access changes them. This causes the familiar loop: the merchant fixes the file, refetches it, sees the block is still there, and assumes the edit failed.
How to check your store
Start with your served robots file. Open a private browser window and fetch https://yourstore.com/robots.txt, with your own domain in place of yourstore.com. Read the whole file. Look for a Cloudflare managed section, marked by comment lines that read # BEGIN Cloudflare Managed content at the top and # END Cloudflare Managed Content at the bottom. If exact case fails, search for the words Cloudflare Managed.
Inside that section, look for User-agent lines naming AI crawlers, followed by Disallow lines. As an example, a managed block may list GPTBot or ClaudeBot with a Disallow for the root path, while your own rules below list the same names with Allow. When both appear, the managed block above is honored first by many readers, and the edge setting can stop the fetch regardless. That pattern means your store is affected even though your own part looks open.
# BEGIN Cloudflare Managed content
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
# END Cloudflare Managed Content
User-agent: *
Allow: /Check each hostname you serve. If the store answers on the apex and on www, fetch robots.txt on both. If you run regional subdomains, fetch each one. Cloudflare settings apply per zone, so one hostname can be clean while another carries the block. Save a copy of each file with the date, so you can compare before and after the fix.
How to read the audit finding
The free Agent Ready audit detects the Cloudflare managed marker automatically. Run the audit on your store domain and open the crawler access section. If the marker is present, the report names it and lists which crawlers the managed block restricts. It reads the served file, not your origin file, so it sees the same combination an outside crawler sees.
Read the finding in three parts. First, the marker line says whether Cloudflare injection is present. Second, the crawler list says which agents the managed section restricts. Third, the guidance says which dashboard setting to change. Run once before the fix and save the link, then run again after. If the second run still shows the block, the change did not apply to the zone you tested, or a cached copy is still served.
How to fix it in Cloudflare
You need dashboard access for the zone that serves your store. Select the domain, then find the AI crawl control area. Look under AI Crawl Control or under Scrape Shield, depending on your dashboard version. You want the Search, Training, and Agent switches, plus a separate robots management dropdown.
- Confirm the zone at the top of the dashboard matches the store domain shoppers use. Repeat per zone if you manage several.
- Keep Search on allow. Search indexing feeds product discovery. Nothing here requires touching it.
- Decide Training deliberately. Keep it blocked if you want your text out of training sets, or permit it if maximum visibility matters more.
- Set Agent to allow if you want shopping assistants to reach your pages. This switch controls the fetchers behind live shopping answers.
- Set the "Manage your robots.txt" dropdown to disable. This stops the managed block injection above your own rules.
- Save, refetch robots.txt in a private window, and rerun the audit. Confirm the marker lines are gone.
The dropdown step is the one most owners miss. Flipping the category switches without touching it leaves the managed section in the served file. Setting it to disable removes the injection, so your own file becomes the served file again. Verify by refetching and confirming the BEGIN and END lines are gone.
Do not confuse the two choices. Blocking training crawlers is a data use choice. Blocking shopping agents is a sales channel choice. Keep Search allowed, set Agent per your shopping goal, set Training per your data choice, and turn the injection off. Allow a day for cached copies to clear before judging crawler behavior in logs.
On Shopify
Cloudflare often sits in front of Shopify custom domains. The merchant points DNS to Cloudflare, then manages the store in Shopify as usual. Shopify serves the pages and the sitemap, but robots requests pass through Cloudflare first. That is enough for the managed block to appear in the served file.
Where to check on Shopify specifically. First, fetch robots.txt on the custom domain shoppers use, not on the myshopify.com address. Second, open Online Store, Themes, the theme menu, Edit code, Templates, and open robots.txt.liquid to see your own rules. If the served file holds a managed section that the template lacks, the extra part comes from Cloudflare. Do not try to remove it in the template. The template cannot remove edge injection.
Who can fix it. The person or agency that controls the Cloudflare zone for your custom domain must make the change. Ask them to confirm the zone, the Agent setting, and the robots dropdown state, and to send before and after copies of the served file. Keep that record with your theme notes, and recheck the served file after every theme update or domain change.
What not to do
Do not open everything indiscriminately. Permit the shopping agents you want and keep the rest of your posture. Keep firewall rules, rate limits, and checkout and account Disallow rules. As an example, a store might permit Agent fetchers while keeping Training blocked and keeping strict limits on unknown scrapers. That is a deliberate split, not a blanket open.
Do not assume a clean robots file means a clean edge. After any dashboard change, and once per month in any case, refetch the served file and look for the marker lines. Plan changes, zone moves, and preset profiles can reintroduce the managed section without warning. A one minute fetch guards against a silent reblock.
Do not copy a starter template that blocks AI crawlers by name without reading it. Many still ship with Disallow lines for GPTBot, ClaudeBot, or Common Crawl. On a store that wants AI shopping traffic, those lines hide the catalog from the assistants you want. Read every named block, keep the internal Disallow rules you need, and permit the answer fetchers by name.
Sources
Four pages informed this guide. The Bytevyte report on the September 15 block describes the default change and notes that the Agent category, which retrieves pages for AI assistants and agents, is blocked by default. The Atwix ecommerce agency note on the Agent category advises merchants to think twice before blocking Agent bots, since those fetchers act for a buyer researching a purchase, and notes that managed injection can list ClaudeBot and GPTBot above merchant rules. The walkthrough of the Control AI crawlers panel shows where the switches live and explains that the Manage your robots.txt dropdown causes the managed block and must be set to disable. The Dev to piece on Search, Training, and Agent separation explains how Cloudflare replaced one broad AI bot switch with separate controls per category.
Panels and labels change over time. If the names above do not match your dashboard, search Cloudflare docs for AI Crawl Control and Scrape Shield and follow the current layout. The idea stays the same: Search, Training, and Agent are separate choices, the managed robots option controls injection, and the served robots file is the evidence.
Want to know if the Cloudflare block applies to your store today. Run the free Agent Ready audit to get a graded score and the exact fixes that matter most for your store. It takes under a minute and needs no signup.
Get a free AI-crawler audit of your store
The audit is free and needs no signup. Enter your email and store address and we will also email you when your score changes.
More guides: ChatGPT Shopping Shopify schema robots.txt for AI AI-crawler friendly Product schema llms.txt AI bot rules FAQ schema Cloudflare AI block