top of page

AI Crawler Access: What Your Store Tells AI Crawlers

Writer: Jim Boudreau
Jim Boudreau
1 day ago
4 min read

Every conversation about AI search right now is about content. Write better product descriptions. Add structured data. Answer the questions buyers ask.


All reasonable. All completely irrelevant if the machine cannot reach the page in the first place.


Abstract image illustrating a an AI search bot combing through service in search of data

On September 15, 2026, Cloudflare changed its defaults for AI crawler access. The version of that change you may have seen in a headline probably does not apply to your store. Something quieter almost certainly does, and it takes about ten minutes to rule out.


What Actually Changed on September 15th


Cloudflare sits in front of a large share of the web, including a great many Shopify, BigCommerce and WooCommerce storefronts. The company announced a new Content Signals Policy on 1 July, with the defaults taking effect on September 15th.


The policy sorts automated visitors into three classifications. Search and indexing crawlers stay allowed — in Cloudflare's words, "Search will remain allowed by default." Training crawlers, the ones collecting text to train models, are blocked. And the one that matters most to a store is the category in between, what Cloudflare calls agent bots: the crawlers an AI assistant uses in real time when somebody asks it to find, compare or buy something.


Those last two are blocked by default — but read Cloudflare's own scope, because almost nobody has. The new defaults apply "for all new domains onboarding to Cloudflare," and only "on the pages that display ads."


If your store has been behind Cloudflare for a while, nothing reached in and switched your AI crawler access off on the 15th. That is worth saying plainly, because a good deal of the coverage implied otherwise.


Here is what did reach you.


The File You Did Not Write


Cloudflare also introduced a robots.txt extension that describes what a visitor may DO with your content rather than merely whether it may read it — use=immediate for interact but store nothing, use=reference for index and link back, use=full for summarise and reproduce.


And in Cloudflare's own words, "all customers who have already enabled managed robots.txt... will now have the additional preference of use=reference added."


Managed robots.txt is the feature that writes and maintains the file for you. It is genuinely convenient and it is exactly the sort of thing a busy operator switches on gratefully, in about four seconds, some afternoon in a settings panel.


Look at what it actually writes. Per Cloudflare's own developer documentation, the managed file disallows eight named AI crawlers outright — ClaudeBot, Google-Extended, Applebot-Extended, GPTBot, Amazonbot, Bytespider, CCBot and meta-externalagent — each with a flat Disallow: /. It also sets a content signal of search=yes, ai-train=no, use=reference.


That is a substantive position. It may even be the right one for your business; there are real arguments for refusing to let your catalog train somebody else's model for free.


But it is a position most merchants did not knowingly take. They clicked a toggle labelled with the word "managed," which reads like housekeeping, and it quietly wrote a policy on their behalf about who may read their store.


Why This Sits Ahead of Everything Else


I wrote recently about what AI agents actually read when they arrive at your store — the identifiers, the brand field, the description, the parts of the catalog nobody inspects. That work matters.


But it all assumes arrival.


Twenty-eight years in eCommerce have taught me one sequence that never changes, whatever the era. Before you optimize what the machine sees, confirm the machine can see it. We spent the 2000s learning this with robots.txt and noindex tags and staging servers accidentally left blocked, and the lesson did not stop being true because the crawler got smarter. AI crawler access is the 2026 version of a very old mistake.


The failure mode is what makes it dangerous. A blocked crawler does not produce an error in your dashboard. Traffic does not drop, because the traffic never existed. Nothing looks wrong. You simply do not appear, and you have no way of knowing you were eligible to.


How Do I Check My Store's AI Crawler Access?


Ten minutes, free, no tools beyond a browser. Do them in this order.


1. Open your own robots.txt. Type your domain followed by /robots.txt into a browser. Read every line. This single step answers the question for most people and it is the one nobody does.


2. Look for the eight names. ClaudeBot, Google-Extended, Applebot-Extended, GPTBot, Amazonbot, Bytespider, CCBot, meta-externalagent. If you find them with Disallow: / underneath and you did not put them there, managed robots.txt did.


3. Look for a content signal line. search=yes, ai-train=no, use=reference is the managed default. Whatever it says, it is now your public statement about how your content may be used.


4. Find out whether you are behind Cloudflare at all. Many merchants are and do not know it, because a developer set it up years ago or the host includes it. A DNS lookup will tell you, and so will your hosting provider. If you are not, steps 1 to 3 still matter — bad robots.txt rules predate all of this by twenty years.


5. Decide deliberately, and write down what you decided. Refusing training crawlers while permitting the ones that can cite you is a coherent position. So is refusing everything. So is permitting everything. The point is that the position should be yours.


6. Re-check in ninety days. Defaults move. This set moved twice in fifteen months.


What I Would Do Monday Morning


The merchants who will be hurt by this are not the ones who made a bad choice. They are the ones who never knew there was a choice.


That is the whole shape of the thing. Not a penalty, not an algorithm update, not something you did wrong. A file written on your behalf by a feature you switched on for convenience, saying something to the entire automated web that you never said.


Go and read your robots.txt. It will take longer to find the tab than to read the file.


Jim Boudreau has worked in eCommerce for twenty-eight years.

 
 
bottom of page
Our website uses an intelligent AI assistant powered by Ultimo Bots to improve customer service.