Your CDN can overrule robots.txt: finding the layer that refuses AI crawlers

A recent analysis of 92 websites reveals that many site owners are unaware that their servers or CDNs may be blocking AI crawlers despite their robots.txt files allowing them. The investigation found that 24% of tested sites returned a 403 error to AI bots while still serving content to browsers. These blocks often occur at the CDN level (such as Cloudflare) or through server-side configurations like nginx or Apache, rather than via the robots.txt file. The author emphasizes that site owners should audit their infrastructure to ensure that their technical implementation aligns with their intended accessibility policies. The article provides practical debugging steps using curl to identify which layer is responsible for the refusal and offers configuration examples for Cloudflare and nginx to properly manage AI crawler access, ensuring that robots.txt remains the primary source of truth for bot behavior.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
In a new article on Habr, the author examines the issue of ABI incompatibility in the C programming language. An ABI break occurs when different binar…
In a recent Habr article, the author analyzes a critical issue with the chunk() method when processing large datasets in background tasks. Using a pro…
OriginTrace: Protecting the DEV Community from Content Theft using Sanity Context MCP
OriginTrace is a new AI-powered tool designed to help developers identify and combat content theft. Created for the Sanity Challenge, the platform sca…



