Cloudflare AI Crawler Defaults Just Flipped: The Access Decision You Did Not Make
The Cloudflare AI crawler defaults flipped on September 15: Training and Agent bots are now blocked on ad pages. Here is the access decision to make now.

The switch got flipped for you on September 15
Most access decisions on your site are ones you made. The Cloudflare AI crawler defaults that changed on September 15 are the exception: for a large slice of the web, the setting got chosen for you, and most owners never saw the prompt.
Here is the short version. Cloudflare replaced its single block AI bots toggle with three categories, and set new defaults that block two of them on pages that show ads. If your site sits behind Cloudflare and you did not record a choice before the fifteenth, you are now living with whatever the platform decided was reasonable. That is fine if it matches your strategy. It is a problem if you never had one.
I have spent enough years in technical SEO to know the most dangerous config is not a wrong one. It is the one nobody chose. This is that, at internet scale.
What actually changed, in plain terms
Cloudflare split AI crawler traffic into three buckets instead of treating every bot as one blob:
- Search: crawlers that index a page now to answer questions about it later. This is the traditional deal, crawl in exchange for referral traffic.
- Agent: automated systems fetching pages in real time on behalf of a person waiting for an answer. Think of the bots that go get a page the moment a user asks.
- Training: crawlers that pull your content into a model's weights.
The controls went live July 1 for every customer, including the free tier. On September 15, the defaults changed: on pages that display ads, Training and Agent are blocked by default, while Search stays allowed.
Who inherited the new defaults matters. The change applies to new domains onboarding to Cloudflare, new sites added by existing customers, and all existing free-tier sites. If you are a paying customer with settings you configured, you likely kept them. If you are on the free plan and never touched bot management, your posture moved without a deploy, a ticket, or an email you remember reading.
One more nuance that trips people up. Cloudflare's stated goal is to push mixed-use crawlers to declare whether they are doing search, agent, or training work. Crawlers that blend all three and refuse to separate are the ones most likely to get caught by the block. That is the leverage play behind the policy, dressed up as a default.
This is a visibility decision wearing a security costume
Blocking bots feels like hygiene. It is not. On the modern web, those categories map directly to whether you show up in the places buyers now look.
- Block Search crawlers and you fall out of the index. Nobody sane is doing that, and the default leaves it on.
- Block Training and you are betting that your content's future recall inside models is worth less than the principle of not feeding them for free. Reasonable people disagree here.
- Block Agent and you may vanish from the exact real-time fetches that increasingly complete tasks for users. That is the one to think hardest about, because agentic retrieval is where a growing share of high-intent activity is heading. I wrote about what that shift does to your funnel in what the agentic commerce protocol does to your feed, and the crawler question is the front door to all of it.
Here is the honest part. If your site does not run ads, the default barely touches you, and there is little upside to blocking crawlers that could cost you discoverability. If you monetize with ads, the calculus is real: you are weighing referral value and AI visibility against the cost of your work being used without a click. That is a business decision. Do not let a CDN default make it silently.
The Crawler Access Decision Grid
Use this to turn a vague "should we block AI?" into three specific calls. For each category, answer the business question, then set the control to match.
1. Search crawlers. Question: do I want to be found at all? Answer is almost always yes. Allow. This is the same crawl-for-traffic bargain SEO has always run on. Confirm it is allowed and move on.
2. Agent crawlers. Question: do I want to be the source an assistant fetches when a user is mid-task? For most brands with commercial intent pages, yes. Allow, and treat agent readability as a ranking-adjacent priority. If you sell something or book something, being absent from the real-time fetch is worse than being read.
3. Training crawlers. Question: is future model recall worth more to me than the principle of paid use? This is the only genuine toss-up. Allow if brand recall inside answer engines is strategic to you. Block if you are a premium publisher whose content is the product and you intend to negotiate. Meter if a compensation path exists and you want to test it.
Run every page type through the grid, not just the homepage. Your ad-supported blog, your gated resources, and your product pages can each deserve a different answer.
The 20-minute crawler audit
Before you accept the inherited defaults, do this:
- Log into your CDN and find the AI bot policy. Note whether you are on a plan that got the new defaults.
- Record your current posture per category. If it is blank, it is not blank anymore, it is the default.
- Probe your own site with an AI user agent and confirm what actually returns. Do not trust the dashboard label, trust the response.
- Separate ad pages from non-ad pages, since the default only bites where ads render.
- Check that your robots directives and your edge rules agree. They often do not.
That last point deserves emphasis. Robots.txt is a request, not a wall. Compliance with it is voluntary, so it states a preference and stops nothing at the technical level. Edge rules enforce. If you have been relying on robots.txt alone, you have been posting a sign, not locking a door. I broke down the difference in the art of crawl control, and the newer file layer in controlling how AI crawlers use your site with llms.txt. Read both against your edge config, because the CDN default now overrides polite intent.
The compensation question is still open
The reason any of this is happening is money. Cloudflare is evolving its pay-per-crawl idea into a usage-based model meant to pay publishers when their content is actually used in AI products, not merely fetched. A couple of AI companies are piloting it. It is early, availability is limited, and it does not yet answer the hardest case: what a publisher is owed when work is used to train a model but never surfaces in a cited answer.
My read: treat metering as an experiment to watch, not a plan to bank on. The durable move is still to be visible where decisions get made and to earn the brand demand that survives any click economy. Blocking everything to protest the terms mostly hurts you. Feeding everything without a thought is how you become a commodity input. The middle is a deliberate posture, set per category, reviewed each quarter. If you want the strategy side of staying visible while models answer, start with generative engine optimization.
The move this week
The deadline already passed. That does not mean the decision is closed, it means the clock is now measuring how long you run on a setting you did not pick. Pull up your CDN, run the grid, record a choice per category, and make it match your actual strategy instead of a platform's opinion of it.
Numbers over noise: the sites that win the next year are not the ones that blocked hardest or opened widest. They are the ones that knew exactly which door they left open, and why.
If your team wants a second set of eyes on your crawler posture before the next quarterly review, the channel is open by introduction. Bring your config, and we will map it to the grid.
Written by Joseph Carroll, Carroll Consulting Services. Connect on LinkedIn ↗
