A physical railway switch lever on a dark background with orange edge lighting, representing Cloudflare routing AI crawler traffic down different tracks.
All Posts

Cloudflare's September 15 AI Crawler Defaults: What Changed for Agents and Data Teams

Franklin Uche
Franklin Uche · Community Lead
Open markdown

On September 15, 2026, Cloudflare changed what "blocked" means for AI traffic. New domains that earn money from ads now start with AI Training crawlers disallowed and AI Agents blocked on ad-bearing pages, while Search stays open. Cloudflare says more than 20% of the web sits behind its network (Cloudflare Blog, 2026), so a default here works a lot like a policy for a fifth of the web. Two weeks in, here's what actually changed, who it touches, and what teams running AI agents or data pipelines should do about it.

Key Takeaways
  • Since September 15, 2026, Cloudflare sorts AI traffic into three buckets (Search, Training, Agent), and ad-monetized new domains start with AI Training disallowed and Agent traffic blocked on pages with ads (Cloudflare Blog, 2026).
  • In June 2026, 52% of crawler requests Cloudflare saw were for AI training, up from 22% in spring 2025 (Cloudflare Blog, 2026).
  • If your agent fetches pages on behalf of a user, you're now in the "Agent" bucket. Plan for per-page variance, not per-site rules.

What exactly changed on September 15?

On September 15, 2026, Cloudflare changed how its bot controls handle AI traffic (Cloudflare Blog, "Have it both ways", 2026). Its "Block" and "Block on pages with ads" settings now also apply to mixed-use crawlers, the bots that collect pages for both search and AI. The single "Block AI Bots" toggle is being retired, and a new "Disallow AI Training" setting lets a site stay in search results while refusing training. Managed robots.txt also becomes "Bot Preference Sync." The practical change is the split. Cloudflare used to treat "AI bots" as one category. It now sorts AI traffic into three buckets, Search, Training, and Agent, and gives each its own switch. New domains are offered one of two presets, depending on whether the site earns money from advertising:

Bucket What it covers New ad-monetized domain New non-ad domain
Search Indexing for search results Allowed Allowed
Training Collecting content to train models Disallow AI Training Allowed
Agent Fetching pages for a user's live request Blocked on pages with ads Allowed

Presets as listed in Cloudflare's September 15, 2026 post, Have it both ways.

So the web didn't flip overnight. Existing customers' settings migrate automatically, per Cloudflare, and the new presets apply only when a customer onboards a new domain. What changed is the starting point for every new site from here on, and on a network that size, starting points add up fast.

Why is Cloudflare splitting crawlers into Search, Training, and Agent?

Cloudflare split its AI controls because training traffic has overtaken search traffic. As of June 2026, 52% of crawler requests Cloudflare observed were for AI training, up from 22% in spring 2025, and mixed-use crawlers made up over 36% of activity (Cloudflare, "Content Independence Day, one year on", 2026). Pure search crawling is now a small and shrinking share. For a site owner, that means most of the AI bots reaching their pages are collecting content for models rather than sending readers back through search results, and a single "AI bots" switch had no way to tell those apart.

Publishers feel that as a lopsided trade. TollBit, which helps publishers license content to AI companies, tracks this in its State of the Bots report. The Q1 and Q2 2026 edition, covering AI bots from 40 vendors across 3,906 publishers, showed the scrape-to-referral ratio climbing from 150:1 in Q1 to 227:1 in Q2 (TollBit, 2026, as reported by Digiday in August 2026). That's hundreds of bot visits for every human who clicks through.

So why not just block everything? Because less than 1% of Cloudflare sites block search bots, according to the same September post, while 17% turn on some way to block training (Cloudflare Blog, 2026). Sites want Google traffic. They just don't want to hand over training data for free. The three-bucket split lets them say yes to one and no to the other.

For anyone building AI products that browse on a user's behalf, the "Agent" bucket is the one that matters. Training crawlers were bound to get squeezed. Agents are newer, and they now get their own switch and their own default instead of sharing one "AI bots" toggle with training crawlers.

What is an "accountable" mixed-use crawler?

Cloudflare defines an accountable mixed-use crawler as one whose operator gives site owners a way to opt out of AI training (through robots.txt or a similar standard), opt-out controls for AI summaries, URL-level visibility into training use, and assurance that opting out won't hurt search rankings (Cloudflare Blog, 2026). Cloudflare names Applebot, Bingbot, and Googlebot as accountable mixed-use crawlers. It lists Amazon, Anthropic, Meta, and OpenAI as "accountable separated" operators, because they run distinct crawlers for search and for training. Accountable mixed-use crawlers stay allowed for search on sites that pick "Disallow AI Training." Every other training crawler gets blocked on those sites.

This sets up a clear incentive. If you run a crawler, the path to staying welcome is to separate your purposes or give sites real controls over each one. A single bot that does everything is now the easiest thing to block.

Pay Per Crawl becomes Pay Per Use

Under Pay Per Crawl, Cloudflare's earlier model, sites could charge AI crawlers a fee for each page request. In July 2026, Cloudflare said it would pay publishers when their content shapes an AI answer, not only when a page gets fetched. The AI search companies Ceramic.ai and You.com were named as early partners (The Next Web, 2026). That's a real shift in what gets priced. A crawl is a single request, while a "use" is the moment content ends up in an answer someone reads. For publishers, it ties payment closer to value. For AI companies, it means the cost of content could start to track usage instead of volume, which changes the math on whether to fetch a page at all.

We covered the earlier version of this model, and why the open web started closing to bots, in the closing web: AI crawler blocking, pay-per-crawl, and what it means for agents.

What should teams running AI agents do now?

Start by accepting that access now varies page by page. Under the new ad-monetized preset, an agent can reach a site's pricing page and then get blocked on the ad-supported article next to it.

A practical checklist:

  1. Know which bucket you're in. Fetching a page because a user asked a question is Agent traffic. Collecting pages to fine-tune a model is Training. Label them separately in your own logs, because the web now treats them differently.
  2. Expect per-page answers. "Block on pages with ads" means the same domain can say yes and no. Build retries and fallbacks per URL, not per domain.
  3. Treat blocks as signals. A block under these presets is a publisher's stated preference. Log it and route around it with a licensed or structured source where one exists.
  4. Watch Pay Per Use. If a source you depend on joins a paid program, paying may be cheaper than engineering around it.
  5. Separate your workloads, and identify them. If you run both training and agent traffic, keep them on different identities so one doesn't get the other blocked. Cloudflare also runs a "signed agents" program: agents directed by an end user can sign their HTTP requests with Web Bot Auth, a cryptographic message-signature standard, so Cloudflare can verify who they are (Cloudflare Docs, Signed agents). If you operate an agent at any scale, it's worth reading.

One trap is worth calling out. For agent pipelines, the failures that hurt most often aren't hard blocks, which are easy to see. Cloudflare's own Challenge Pages are the easy case: every one carries a cf-mitigated: challenge header, and its content type is always text/html, whatever resource you asked for (Cloudflare Docs, Detect a Challenge Page response). Test for that header before you parse anything. The harder case is a silent partial page: a fetch returns a 200 status, but the body is a placeholder page (an interstitial), a JavaScript shell, or a stripped version of the article. Check what came back, not just the status code. A minimum content length, or a check for a DOM element you expect on the real page, catches most of these.

How Massive fits Cloudflare's new defaults

Massive provides a device-access network and rendering stack that delivers clean HTML or markdown from public pages, in any of 195+ countries. Customers run their own operations on top. Publishers' AI preferences are public signals, and the teams we work with treat them as inputs to their own policy, not obstacles to route around.

The Web Render API returns rendered pages as markdown for LLM pipelines, and lets customers geotarget requests by country, region, or city. On Residential Proxies, each account also keeps its own domain blocklist, so a team can stop requests to any site it has decided not to access. The list holds up to 1,000 domains, and a request to a blocked domain is refused with a 452 (Disallowed Content) error before it ever reaches the site. If your agent stack is hitting the new defaults, see why AI agents get blocked on datacenter IPs and how to fix it, or step back to the full guide on how to give AI agents live web access. For the legal side of training data, read the AI training data lawsuits every web data team should know.

The bottom line

  • Cloudflare now treats Search, Training, and Agent traffic as three separate decisions, each with its own switch and its own default for new domains.
  • Ad-monetized new domains start with Training disallowed and Agents blocked on ad pages.
  • Pay Per Use pays publishers for answers, not fetches.
  • Agent builders should plan for per-page access, check every response for the cf-mitigated: challenge header, and log each block as a publisher's stated preference.
  • Next to watch: which AI companies join Pay Per Use, and how many agents register as signed agents.

Want to see how rendered, geotargeted fetches look in practice? Start with the Web Render API overview.

Sources

Frequently Asked Questions

Did Cloudflare block all AI crawlers on September 15, 2026?+

No. The new defaults apply to new domains onboarding to Cloudflare, and existing settings migrated automatically. Search stays allowed by default everywhere. Training and Agent traffic are restricted by default only on ad-monetized new domains, according to Cloudflare's September 15, 2026 post.

Is Googlebot blocked by the new Cloudflare settings?+

Not on sites that use "Disallow AI Training." Cloudflare classes Googlebot, Bingbot, and Applebot as accountable mixed-use crawlers, which stay allowed for search. Sites that choose a full "Block" setting, though, now block mixed-use crawlers too, which can affect search visibility.

What's the difference between Pay Per Crawl and Pay Per Use?+

Pay Per Crawl charged AI crawlers per request. Pay Per Use, announced in July 2026, pays publishers when their content shapes an AI answer. Ceramic.ai and You.com were named as early partners, per The Next Web's coverage.

How much crawler traffic is AI training now?+

As of June 2026, Cloudflare reported that 52% of crawler requests on its network were for AI training, up from 22% in spring 2025. Mixed-use crawlers accounted for over 36% of activity.