On September 15, Cloudflare begins blocking AI crawlers by default for new sites. It reads like an on/off toggle for AI. It is really three separate decisions, only some are made for you, and the labels that drive them are written by the AI companies themselves. Read it as one switch and you can block the wrong thing, keep the thing you meant to stop, and lose Google along the way.
Cloudflare sits in front of roughly a fifth of the web, and from September 15 it blocks AI crawlers by default for new domains onboarding to it (Cloudflare's own announcement; TechCrunch coverage). Existing owners keep their settings, and anyone can opt out before the date. It arrives alongside Pay-Per-Crawl, where a publisher can answer a crawler with a price rather than a wall. So the block is also a paywall and a bargaining position, not just a gate.
Underneath, Cloudflare sorts crawlers by declared purpose, and the difference between them is the whole point:
The default blocks Training and Agent on pages that display ads — the monetizable pages a site owner meant a person to land on — and leaves Search allowed. Three different value exchanges, collapsed by most people into a single "block AI" reflex.
Because Search stays on by default, flipping the switch does not remove you from the crawlers that feed AI Overviews and assistant answers. If your goal was to stop showing up in AI results, the default does not do it. If your goal was to stop training without losing AI visibility, the default is closer to right than you would guess. Either way, you cannot tell which happened without looking at what actually reached you.
Here is the trap almost no one reads to the end of. The big crawlers are multi-purpose: Googlebot, Applebot and Bingbot crawl for ordinary search and for AI at the same time. Cloudflare enforces the most restrictive rule that matches, so the moment you choose to block Training, those mixed crawlers are blocked too — unless you specifically opt out for "training crawlers that also crawl for search." Reach for the training block without that exception and you can dent your normal Google and Bing visibility while trying to keep your work out of a model. The blunt version of this decision has collateral damage the label never warns you about.
Every one of these categories rests on a claim. Purpose is self-reported, and the user-agent that carries it is the one part of a request the sender writes themselves. A category built on an unverified identity inherits that uncertainty: in the crawler traffic we have verified, more than one in five requests wearing an AI bot's name could not be authenticated against the operator's published ranges at all. Which is the other half of this decision, the one we covered in verify before you block. The short version: match the source IP to the operator's published range and forward-confirm it by reverse DNS, because the name alone proves nothing.
This is what we are for. Unsourced shows which AI crawlers reached your pages, verified by identity rather than the label they declare, what each one fetched, and whether that visit became a citation across the assistants we check. An optional edge worker turns away crawlers whose claimed identity it can disprove, lets the verified ones through, and keeps the log — and it fails open, so a genuine crawler is never at risk. You can also test any single request against published ranges and forward-confirmed reverse DNS with our free crawler verifier. Handled that way, September 15 is a decision you make with evidence, not a default you inherit.
The block is a real, useful tool. It is just not a light switch. It is three decisions keyed on labels you did not write, one of which can take Google down with it. Publishers are being handed switches from every direction now — Google shipped a "Preferred Sources" button in August for the same anxiety. But a switch without measurement is a guess with a nicer interface. See which bots reach you, and what they are, before you flip anything.
For new domains onboarding to Cloudflare, crawlers classed as Training or Agent are blocked by default on pages that display ads, while Search stays allowed. Existing owners keep their current settings, and anyone can opt out of the change in their security settings before the date.
Training crawlers absorb your pages into a model's weights, with no referral back. Search crawlers index your pages so an assistant can retrieve and cite them at answer time. Agent crawlers fetch a page in real time because someone's AI assistant is acting for them right now. Cloudflare allows Search by default because it can funnel visitors back, and blocks Training and Agent because they typically do not.
It can. Multi-purpose crawlers such as Googlebot, Applebot and Bingbot crawl for search and for AI at once. Cloudflare enforces the most restrictive matching rule, so choosing to block Training also blocks those crawlers unless you opt out for training crawlers that also crawl for search. Blocking model training without meaning to dent ordinary search visibility takes a deliberate setting, not the default.
Not by itself. Because Search-purpose crawlers stay allowed, the pipe that feeds AI Overviews and assistant answers keeps running. If your aim was to leave AI results entirely, the default does not do it. If your aim was to stop training while staying visible, the default is closer to right than most people assume.
Treat them as claims. Purpose is self-reported and a user-agent can be spoofed, so a category is only as trustworthy as the identity behind it. Verify the source IP against the operator's published ranges and forward-confirmed reverse DNS before you rely on the label. In our own traffic, more than one in five requests wearing an AI crawler's name failed that check.
The other half of this decision: the user-agent you would block on is a claim, and more than 1 in 5 are impostors.
Read it →Every major AI crawler: who it is, how to verify it, and whether to allow or block it.
Read it →Our 30-day data: 21.8% of AI-crawler hits failed verification.
Read it →Unsourced captures which AI assistants cite you, proves which crawlers really fetched your pages, and re-checks after you act — evidence, not a score.
© Unsourced, the evidence layer for AI search. Our data and findings are free to quote and cite — please attribute to Unsourced and link to unsourced.app.