Cloudflare launched Bot Preference Sync , a tool that automatically generates and updates your site's robots.txt file based on rules you set for AI crawler categories — not bot by bot. You can decide, for example, to allow OpenAI crawlers and block ByteDance ones with a single category choice, without writing a line of code.
For those managing a business website, an e-commerce store, or an industry blog, this makes a real difference: visibility in AI results (ChatGPT, Perplexity, Google AI Overviews) also depends on which crawlers you allow through. Until now, keeping that list updated was manual work and often overlooked. Starting today, Cloudflare can do it for you.
If your site already uses Cloudflare as a CDN or firewall, it's worth exploring this feature this week. If you don't use it, the article still explains the logic to understand how to update robots.txt manually with criteria.
The problem nobody wants to face: robots.txt is stuck in 2019
The robots.txt file is a list of instructions that tells bots — search engines, AI crawlers, various spiders — what they can read from your site and what they can't. It's been around for decades, works well, but requires manual updates every time a new crawler enters the market.
In 2026, AI crawlers are multiplying every quarter. GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot, Bytespider (ByteDance), CCBot (Common Crawl) — and more are coming. Each crawler has a different name in robots.txt. Who keeps that list updated? Almost no one.
The typical result: a robots.txt configured years ago, AI crawlers entering or exiting randomly, and no real strategy on who can index your content for AI responses. If you're trying to understand how the AI visibility as a metric for SMEs , robots.txt is one of the first places where you lose ground without realizing it.
What Cloudflare's Bot Preference Sync does
Cloudflare presented Bot Preference Sync : a tool that automatically generates and updates your domain's robots.txt, based on rules you set for Category of crawlers, not per single bot.
The logic is this: instead of writing one line for GPTBot, one for ClaudeBot, one for PerplexityBot, and so on, you choose a category — 'AI crawlers for training,' 'AI crawlers for real-time responses,' 'traditional search engine crawlers' — and Cloudflare translates that choice into specific rules, updated as new bots arrive in its database.
As reported by Search Engine Journal , the main point is that no single setting fully describes what a company wants to do: allowing GPTBot and blocking Bytespider are two different choices, with different reasons, and a category-based tool makes them manageable without advanced technical skills.
Three steps to use it — or do it manually if you don't use Cloudflare
If your site is already on Cloudflare:
- Log in to the domain's Cloudflare dashboard.
- Look for the Bot Management section (available in Pro plans and above) and check if Bot Preference Sync is already active or rolling out for your account.
- Set preferences by category: decide which AI crawlers you want to allow and which to block, based on how they use your content.
If you don't use Cloudflare:
- Open your robots.txt (usually at yourdomain.com/robots.txt).
- Manually add directives for the main AI crawlers: GPTBot, ClaudeBot, PerplexityBot, Bytespider, CCBot. Each crawler has a specific name to use in the line
User-agent. - Decide for each one if you want
Allow: /orDisallow: /, based on your AI visibility strategy.
The difference between the two approaches is maintenance: with Cloudflare, the list updates itself. Manually, you have to keep an eye on it every time a new relevant crawler arrives.
Who really needs this feature — and who can wait
It's immediately useful for those who:
- publishes original content (guides, catalogs, industry articles) that it wants to see cited in AI responses;
- has already worked on a content validation workflow for AI Search and wants the right crawlers to reach them;
- manages an e-commerce site or a site with sensitive data on which it doesn't want training by certain models.
Can wait for those who:
- has a showcase website with few updates and no active content strategy;
- doesn't use Cloudflare yet and doesn't have resources for migration right now.
Warning: blocking all AI crawlers out of fear is not a strategy. It means exiting ChatGPT, Perplexity, and Google AI Overviews results. If the visibility in AI responses is already a goal in your marketing plan, blocking the crawlers that feed it is self-defeating.
The typical mistake and what this tool doesn't do
The most common mistake is treating robots.txt like an on/off switch: you either block everything or let everything pass. Cloudflare's category-based logic encourages more granular thinking — and that's the real value of the feature.
What Bot Preference Sync doesn't do :
- It doesn't guarantee that crawlers will respect the instructions. robots.txt is a voluntary protocol: legitimate bots follow it, malicious ones don't.
- It doesn't replace a content strategy. Opening the door to AI crawlers doesn't mean being cited: your content must be authoritative, structured, and relevant. On this, it's worth reading how it works local visibility in AI Search for Italian SMEs .
- It doesn't solve the problem of conflicting brand information already circulating in AI datasets.
robots.txt is one piece of a larger strategy for SEO and AI visibility . Keeping it updated is necessary, but not sufficient.
Related articles
Discover more articles exploring similar topics, selected to offer you a more complete and stimulating perspective. Each piece of content is carefully chosen to enrich your experience.