Skip to main content
Piloterr
Back to blog
September 29, 2026

Cloudflare tightened Bot Management in September 2026. Collectors felt it.

TL;DR: On 15 September 2026, Cloudflare announced Disallow AI Training and new recommended defaults that treat Search, Training, and Agent traffic separately. That update lands on top of Precursor (session-level behavioral detection, rolled out in July) and the usual Managed Challenges / Turnstile stack. If your scrapes or monitors stalled this month, you are not imagining it.

Can the job still finish? That is the only question that matters for most data teams. Piloterr already runs this class of collection through WebUnlocker and the other website products.

What Cloudflare actually shipped in September

The dated announcement is the 15 September Bot Management / AI Crawl Control update, covered in Cloudflare's blog post and press release.

What they published:

  • A new Disallow AI Training setting so a site can refuse training while staying in search (for crawlers Cloudflare labels Accountable).
  • Clearer meaning for Block and Block on pages with ads: those actions now also apply to mixed-use crawlers such as Googlebot, Bingbot, and Applebot. If you only want to stop training and keep search, use Disallow AI Training instead of Block.
  • Deprecation of the blunt Block AI Bots switch in favor of separate Search, Training, and Agent controls.
  • Bot Preference Sync replacing Managed Robots.txt.
  • For new domains onboarding from that date, recommended presets on ad-supported sites: Search allowed, Training set to Disallow AI Training, Agents blocked on pages with ads. The July changelog already previewed those defaults.

Existing zones were migrated according to prior settings. Plenty of operators still flipped options manually after the news cycle. Either way, more domains now classify chat fetchers, browser-use agents, and training-style crawlers as unwanted by default.

That is an official product change. Separately, what many collectors observe in the browser is the older challenge stack getting stricter in practice because more zones turn features on.

Why jobs that used to pass now fail

Cloudflare's own changelog and Precursor docs describe a shift that started earlier in the summer and is still landing on customer zones: detection across the whole session, not a single gate at the door.

Precursor injects client-side JavaScript, collects behavioral signals over time, and updates session state in the cf_clearance cookie. Cloudflare states that clearance can be reduced or invalidated mid-session, and that additional challenges can fire after a visitor already "passed" once. A page refresh does not wipe that session signature the way a one-shot CAPTCHA used to.

So a monitor that looks human for the first HTML fetch can still get challenged on the third pagination hop. A price scraper that clears an interstitial and then hammers product URLs in a tight loop starts looking like agent traffic under the new Search / Training / Agent taxonomy.

None of that requires you to be a named AI crawler. Site owners toggle Agent and Training controls because they saw the September defaults. Your legitimate B2B enrichment job sits in the same automated bucket unless the publisher has carved out an exception.

"HTTP looks fine" but you are not on the real page

This is where pipelines go quiet for the wrong reason.

Cloudflare documents that a full interstitial Challenge Page sets the response header cf-mitigated: challenge, serves content-type: text/html, and (for Managed Challenge / Under Attack Mode) returns a 403 with challenge HTML rather than your product payload. See Detect a Challenge Page response and error page types.

In the field, teams still say "HTTP 200 but stuck on a challenge." Often they mean a softer failure than the documented interstitial.

Some collectors only check status codes and treat any HTML body as success. Challenge HTML is still HTML. If your check is status < 400, you can store "Checking your browser…" as a price.

JavaScript Detections and Precursor are different again: Cloudflare injects them into HTML responses without pausing the visitor the way an interstitial does. You get a real-looking 200 while the edge keeps scoring the session, then fires a challenge later or drops the bot score. Embedded Turnstile on a form that already returned 200 works the same way from an operator's seat. The page loads. The protected action does not finish until the widget clears.

Managed Challenges themselves stay adaptive. Many humans see a short non-interactive wait; traffic that looks automated gets asked to interact. That behavior is in the Challenge Pages docs, not a September-only release note. September pushed more policies at Agent and Training traffic. The challenge machinery is what those policies call.

What operators should expect through autumn

Not every zone flipped on the same day. Still, if you run collectors against Cloudflare properties, plan for more interactive Managed Challenges when fingerprint and behavior signals look off, and for clearance that does not last the whole job once Precursor is on Maximize Security or the session drifts.

Soft failures stay common: responses that look successful until you check cf-mitigated, body length, or whether the payload actually contains the fields you extract. Ad-supported sites that took the new onboarding presets will treat more of what Cloudflare calls Agent traffic as unwanted. Homegrown stacks that "solved Cloudflare" in 2024 or early 2025 often need another maintenance cycle. Session scoring and agent defaults change the failure mode faster than TLS fingerprint tweaks alone can fix.

If you own the collectors, budget for retries that still return challenge HTML, and for validation that looks at content rather than status alone.

If the pipeline is stalling on Cloudflare

Most teams we talk to want the data out, not a second full-time anti-bot project. Piloterr already runs collectors against Cloudflare-protected targets through Website WebUnlocker (allowlisted domains, HTTP mode) and Website Rendering when the page itself needs a real browser. For which tool fits which target: Crawler vs Rendering vs WebUnlocker.

September's Cloudflare changes are real. Absorbing this kind of vendor shift is part of what we do so your product roadmap does not have to.

Need a domain reviewed after the mid-September defaults? Start with Website Rendering when the page needs a real browser, or the WebUnlocker API docs for allowlisted HTTP targets. Open a conversation from the site if you want us to say whether the target is in scope and what success rate to expect.

More to read

Guides and news about web scraping, proxies, and data extraction.

News

Piloterr MCP is on the 2026-07-28 spec

Our hosted MCP at mcp.piloterr.com now follows MCP 2026-07-28: a stateless request/response core, the same Streamable HTTP URL, and the same x-api-key.

Josselin Liebe
Josselin Liebe
Read
News

Understanding p50, p75, p90, p95, and p99 latency metrics

Latency percentiles explain how fast your API or scraping pipeline really performs for most requests and for the slow tail. Learn what p50 through p99 mean, why averages lie, and how to set realistic SLAs.

Josselin Liebe
Josselin Liebe
Read
News

Cloudflare teams with Chrome, Firefox, and Edge on PACT, a privacy-first anti-bot protocol

Cloudflare joins Mozilla, Google, Microsoft, and Shopify to develop PACT (Private Access Control Tokens), a standard meant to authenticate human and authorized agent traffic without CAPTCHAs or invasive tracking.

Josselin Liebe
Josselin Liebe
Read

Ready to get started?

Your web scraping API is one click away. Start with +500 credits, no infrastructure to set up, no proxies to manage, and no credit card required.

  • +500 credits
  • No credit card required
  • All endpoints included