Insights · Search and AI trends
AI crawlers, robots.txt and llms.txt: what a small business should do

If your website runs behind Cloudflare and someone once switched on "Block AI bots", check it today. Cloudflare said that from 15 September it would apply a training block to crawlers that also do search, and it named Googlebot, Applebot and Bingbot. That is the most expensive AI-crawler mistake a small seller can make this month. The cheapest one is paying someone to write an llms.txt file for Google, which Google says it does not use.
- AI companies now run separate bots for search, for model training and for fetching a page when a user asks, and each can be allowed or blocked on its own in robots.txt.
- Google says llms.txt files neither help nor harm visibility in Google Search, and its Search Console setting, not robots.txt, is how a site leaves AI Overviews and AI Mode.
- Cloudflare said that from 15 September 2026 a zone that blocks AI training also blocks crawlers that combine search and training, naming Googlebot, Applebot and Bingbot.
Three kinds of bots, three different decisions
"AI crawler" now covers three different jobs, and the big providers have split them into separate bots so you can choose. A search bot builds the index an AI assistant cites and links. A training bot collects pages that may be used to train future models. A user-triggered fetcher opens your page because a person just asked the assistant about it.
| Provider | Search and answers | Training | When a user asks |
|---|---|---|---|
| Googlebot (Search, including AI Overviews and AI Mode) | Google-Extended (a robots.txt token only) | Google-Agent and other user-triggered fetchers, which generally ignore robots.txt | |
| OpenAI | OAI-SearchBot | GPTBot | ChatGPT-User; OpenAI says robots.txt rules may not apply |
| Anthropic | Claude-SearchBot | ClaudeBot | Claude-User, which has its own robots.txt token |
The consequences are spelled out. OpenAI says sites that opt out of OAI-SearchBot will not be shown in ChatGPT search answers, apart from navigational links, and that a robots.txt change takes about 24 hours to register. Anthropic says blocking Claude-SearchBot or Claude-User may reduce your site's visibility in Claude's search results (help page updated 7 April 2026). Google says Google-Extended controls training and grounding for Gemini Apps and Vertex AI, and does not affect inclusion or ranking in Google Search (crawler list updated 14 July 2026).
For your company: if you sell to people who might ask an assistant about you, allow the search bots. Training is a business decision: each company documents GPTBot, ClaudeBot and Google-Extended separately from its search bot, and OpenAI and Google state that blocking the training bot does not affect search. Blocking user-triggered fetchers mostly means the assistant cannot read your page at the moment a customer asks about it.
Google: robots.txt cannot split Search from AI answers
Googlebot crawls for Search as a whole, and AI Overviews and AI Mode are part of Search. If you want to stay in normal results but leave the AI features, robots.txt is the wrong tool. Google's tool is the Search generative AI control in Search Console, under Settings, which reached all websites on 31 August 2026. Choosing "Exclude" removes your links and content from AI Overviews, AI Mode and generative AI features in Discover. Google says you then get no traffic or impressions from those features, other sites' content may still appear there, and the setting is not used as a ranking signal elsewhere in Search. It does not cover training; for that, Google points to Google-Extended.
For your company: for most small sellers the default, "Include", is the right answer. Excluding your site does not stop people asking the question; it only removes you from the answer.
The Cloudflare change that can block Googlebot
On 1 July 2026 Cloudflare announced new controls that sort AI traffic into Search, Agent and Training, available on every plan including Free. It scheduled two changes for 15 September 2026. New domains would get Training and Agent bots blocked by default on pages that show ads, with Search allowed. And crawlers that combine search with training would be judged by the strictest rule that applies. In Cloudflare's words, multi-purpose crawlers "such as Googlebot, Applebot, and BingBot will be blocked" for customers who chose to block Training, including through the older one-click "Block AI bots" option.
This matters more than robots.txt, because robots.txt is only a request. The standard that defines it (RFC 9309, September 2022) says its rules are not a form of access authorization. A block at your CDN or firewall is enforced. On 4 August 2026 Search Engine Journal reported a site owner's claim that setting AI Training to Block in Cloudflare left Googlebot and Bingbot with 403 errors on the sitemap. That report was not confirmed by Cloudflare or Google, but it shows what to test.
For your company: if you use Cloudflare, open your AI crawler settings and look at the Training line. If it says Block and you still want Google and Bing traffic, find out how your zone now treats Googlebot and Bingbot, and change the setting or add an exception. Then run a live URL Inspection in Search Console on your home page and one product page; a 403 means Googlebot is being turned away. If you use another CDN, a firewall or a security plugin, look for a similar AI-bot switch.
llms.txt: optional, and not for Google
llms.txt was proposed by Jeremy Howard on 3 September 2024: a Markdown file at the root of a site with its name, a short summary and links to key pages, meant to help AI tools at the moment they need information rather than for training. Google added a note to its documentation on 15 June 2026, and its guide to generative AI features (updated 10 July 2026) is direct: you do not need llms.txt or any special AI file or markup, Google Search does not use them, and keeping one for other systems will neither harm nor help your visibility.
Chrome's Lighthouse has an llms.txt check among its agentic browsing audits (page updated 5 May 2026). It flags a server error when the file is requested and marks a missing file as not applicable, because the file is optional. The OpenAI and Anthropic crawler pages cited here describe robots.txt controls and do not mention llms.txt.
A simple default for a small seller. Allow: Googlebot, Bingbot, OAI-SearchBot and Claude-SearchBot. Decide on business grounds: GPTBot, ClaudeBot and Google-Extended. Leave alone: user-triggered fetchers, which act for a real person. Skip: paying for llms.txt "optimisation" aimed at Google. Check: your CDN or firewall, not just robots.txt.
For your company: write an llms.txt only if you already have documentation, a large catalogue or a developer audience, and it takes you less than an hour. If someone offers you a Google ranking gain from it, Google's own guide says otherwise.
A 30-minute check
Open yoursite.com/robots.txt and read it as a stranger would. Search your own robots.txt for GPTBot, ClaudeBot, OAI-SearchBot and Google-Extended, and make sure each line reflects a decision someone actually made. Check the CDN setting above. Then look at the Search Console setting for generative AI and confirm it says what you intend. If you would rather have us look, HB Crossborder's free visibility check reviews your site basics and tests five real customer questions in ChatGPT, Perplexity and Google AI Overviews. See our services, or send one URL and get a free visibility check within 48 hours.
- Open your robots.txt and list every AI bot it names, with the reason each rule is there.
- Log in to your CDN or firewall and check whether AI training is blocked; if it is, test how Googlebot and Bingbot are treated.
- Run a live URL Inspection in Search Console on your home page and one product page and look for a 403 response.
- Confirm that the Search generative AI setting in Search Console says Include, unless you have decided otherwise.
- Allow OAI-SearchBot and Claude-SearchBot if you want to appear in ChatGPT and Claude answers.
- Drop any paid task that promises Google gains from an llms.txt file.
Send us one URL and get a free visibility report within 48 hours.
- Optimizing your website for generative AI features on Google Search, Google Search Central (2026-07-10)
- Latest documentation updates (entry: Clarifying guidance on llms.txt files, 15 June 2026), Google Search Central (2026-06-15)
- Search generative AI control, Google Search Console Help (2026-08-31)
- List of Google's common crawlers, Google Crawling Infrastructure (2026-07-14)
- List of Google user-triggered fetchers, Google Crawling Infrastructure (2026-08-19)
- Overview of OpenAI crawlers (version archived 17 September 2026), OpenAI (2026-09-17)
- Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic (Claude Help Center) (2026-04-07)
- Your site, your rules: new AI traffic options for all customers, Cloudflare (2026-07-01)
- Report That Cloudflare AI Bot Blocking Prevents Googlebot From Indexing Sites, Search Engine Journal (2026-08-04)
- RFC 9309: Robots Exclusion Protocol, IETF (2022-09)
- The /llms.txt file, llmstxt.org (Jeremy Howard) (2024-09-03)
- llms.txt (Lighthouse agentic browsing audit), Chrome for Developers (2026-05-05)
General information, not legal or tax advice. Check with a professional before acting on it.