← Marketing issues

Cloudflare flips AI crawlers to blocked-by-default on September 15

Cloudflare's new crawler policy, announced July 1, actually takes effect on September 15. Free-plan users and new sites will have AI training and agent crawlers blocked by default on ad-supported pages, while search crawlers stay allowed. Do nothing, and this is the setting you get.

Cloudflare, which sits in front of a large share of the web's traffic, rewrote its crawler rules on July 1. It split crawlers into three types — search, AI training, and AI agent — and starting September 15, ad-supported pages will block training and agent crawlers by default, while search crawlers stay allowed. That default applies automatically to free-plan users, new signups, and any new site an existing customer spins up. If you touch nothing, September 15 is the day this setting takes over your site.

Why this is news again now

The announcement is two months old, but the actual switch flips in two weeks — which is why explainer posts telling site owners exactly what to check are picking back up this week. There's also a catch that's easy to miss on first read: crawlers that blend search and training in one, like Googlebot, can get blocked the moment training crawlers are blocked, no matter what you set for search. Trying to protect your search traffic could end up cutting off the very crawler that traffic depends on.

The principle: blocking a crawler and blocking a citation are different calls

What this policy sorts isn't "who's taking my content" but "what they do with it." Search crawlers index your pages for search visibility, training crawlers feed model training, and agent crawlers fetch your page live when a tool like ChatGPT or Perplexity is building an answer. Treat all three as one thing to shut out and you also close the door on your content ever getting cited in an AI answer. Leave all three open and you might just be training data with nothing back. This isn't an on/off switch — it's a decision about which of the three you actually want.

What's confirmed and what isn't — This new default only applies to sites running through Cloudflare's proxy (the "orange cloud"). Sites on hosts that don't route through Cloudflare, like Vercel or Netlify, aren't affected — though plenty of site owners don't actually know whether their setup uses Cloudflare or not. Whether existing paid-plan customers are fully exempt is described a little differently across outlets; we couldn't find one clean official line on it, only that "new sites and free-plan users" are consistently named as the core target. And since the block hasn't taken effect yet, nobody has measured how much AI-citation traffic actually moves once it does.

If your site gets 100 visitors a week

There's one thing to do today: check whether your site runs through Cloudflare (look for a Cloudflare badge in your host's dashboard, or an orange-cloud icon on your DNS records). If it does, go into the security settings before September 15 and decide, deliberately, which of the three crawler types to allow. If getting cited in AI search answers is a goal — as it is for this site — leaving the agent crawler open works in your favor; the training crawler is a separate call to make on its own. The smaller and newer your site is, the more this one decision matters: don't let a default setting decide your content's fate instead of you.

Sources

Issue posts are researched and drafted by an automated AI pipeline, then published through editorial gates: a real event, linked sources, and a stated fact-check date.

Share this postXThreadsLinkedInHacker News

More issues