Guides / How the engines find and choose businesses
llms.txt and AI crawler access: a technical guide for business websites
Which AI crawlers a local business should allow, which it can safely block, how hosts and CDNs shut them out by accident, and what llms.txt does and does not do.
By Ben Hinton
, published
The question: How should a business website handle llms.txt and AI crawler access?
The short answer
Allow the AI search crawlers. Decide separately about the training crawlers. Make sure your host or CDN is not blocking the search crawlers behind your back. Add an llms.txt file if you like, but do not expect it to change who the engines recommend.
The search crawlers are what put your pages in front of ChatGPT, Claude, and Perplexity when they look things up. The training crawlers collect text to train future models. The engines document the two separately, and a local business can block training while staying fully visible in search.
The crawlers that matter, by engine
Each engine names its crawlers in its own documentation.
OpenAI (ChatGPT). OAI-SearchBot1 surfaces websites in ChatGPT's search features, and OpenAI says sites that opt out of it are not shown in ChatGPT search answers, though they can still appear as navigational links. GPTBot2 crawls content that may be used to train OpenAI's models, and it is controlled independently: a site can block GPTBot and allow OAI-SearchBot.
Anthropic (Claude). Anthropic runs a training crawler, ClaudeBot, a bot for fetches that a user asks for, Claude-User, and a search crawler, Claude-SearchBot3. Anthropic says that disabling the search crawler prevents it from indexing your content for search, which may reduce your site's visibility and accuracy in Claude's search results.
Perplexity. PerplexityBot4 surfaces and links websites in Perplexity's search results. Perplexity says it is not used to crawl content for AI foundation models.
Google (AI Overviews, AI Mode). Google's AI features in Search use Google's own index. A page has to be indexed and eligible to appear with a snippet, and Google says there are no additional requirements5 beyond that. In practice that means Googlebot must be able to crawl the pages you care about.
A robots.txt for a local business
robots.txt is a plain text file at the root of your domain that tells crawlers which paths they may fetch. Crawlers read the group for their own user agent name. A sensible starting point for a contractor who wants to be found in AI search but prefers not to be used for training looks like this. The first three groups allow the search crawlers, the next two block the training crawlers, and the last group covers everyone else, including Googlebot:
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: *
Disallow: /admin/
Sitemap: https://www.example.com/sitemap.xml
Whether to block training is your call. Allowing it costs you nothing in search. Blocking it does not remove you from ChatGPT search, according to OpenAI's description of its bots, and Anthropic describes its training and search crawlers as separate bots too. What you must not do is block the search crawlers by accident, and a common accident is a blanket rule. A line like User-agent: * followed by Disallow: /, left over from a staging site or a redesign, shuts out every crawler that has no group of its own.
The block you cannot see in robots.txt
A clean robots.txt is not enough. OpenAI's help center says that to make a website eligible for ChatGPT search, you must allow OAI-Searchbot to crawl the site6, and confirm that the website host or content delivery network allows traffic from OpenAI's published search-bot IP addresses.
That second half is easy to miss. Many website builders, managed hosts, and CDNs offer bot protection that challenges or blocks automated traffic, and a crawler generally cannot get past a challenge page. The robots.txt says yes, and the firewall says no. The owner never sees it, because the site loads fine in a browser.
To check:
- Look at your server or CDN logs for requests from OAI-SearchBot, Claude-SearchBot, and PerplexityBot, and at the status codes they received. A string of
403responses, or challenge pages served with a success code, means the bot is being turned away. - Check your CDN's bot or firewall settings for rules that block AI crawlers as a category, and allow the search crawlers explicitly.
- Where an engine publishes IP ranges for its crawler, as OpenAI does, allow those ranges rather than trusting the user agent string alone, since anyone can claim to be a crawler.
What llms.txt is
llms.txt is a proposed convention, not a formal standard. It is a Markdown file at /llms.txt that gives a short summary of a site and a list of links to its most useful pages, written for language models to read. Some sites also publish llms-full.txt, with the full text of those pages in one file.
It is cheap to make and it does no harm. We publish one on citemy.ai. But be clear about what it is. Google says you do not need to create new machine-readable files, AI text files, or markup to appear in its AI features, and that there is no special schema.org structured data7 to add either. None of the crawler documentation we cite from OpenAI, Anthropic, or Perplexity asks for one. Anyone selling llms.txt as the key to AI visibility is selling a text file.
If you add one, keep it accurate: the business name, what you do, the cities you serve, how to contact you, and links to your service pages. Treat it like your structured data, as one more place where your facts must match everywhere else. See structured data for AI visibility.
What this means for an owner
Crawler access is the one part of AI visibility that is purely technical and entirely in your control. It is also the part most often broken without anyone knowing, because a site can look perfect to people while turning away the bots that matter.
Get the basics right: search crawlers allowed in robots.txt, no blanket block, the firewall and CDN letting them through, and your pages indexed by Google. Then confirm it from the logs, not from the settings screen.
Access only makes you readable. Whether the engines then recommend you depends on what your pages and the sites about you say, covered in how the engines find local businesses. To see what ChatGPT currently says about you, start with how to check whether ChatGPT recommends your business.
Notes
- OpenAI's OAI-SearchBot crawler is what surfaces websites in ChatGPT search; sites that block it are not shown in ChatGPT search answers. OpenAI documentation, checked Sep 27, 2026. Overview of OpenAI Crawlers ↩
- OpenAI's GPTBot crawls content for training its models and is controlled independently of OAI-SearchBot, so a site can block training while still appearing in ChatGPT search. OpenAI documentation, checked Sep 27, 2026. Overview of OpenAI Crawlers ↩
- Anthropic runs three bots - ClaudeBot (training), Claude-User (user-directed fetches) and Claude-SearchBot (search indexing); blocking Claude-SearchBot may reduce a site's visibility in Claude's search results. Anthropic, Apr 7, 2026; checked Sep 27, 2026. Does Anthropic crawl data from the web, and how can site owners block the crawler? ↩
- PerplexityBot is the crawler that surfaces and links websites in Perplexity's search results and is not used for model training; Perplexity recommends allowing it to appear in results. Perplexity documentation, checked Sep 27, 2026. Perplexity Crawlers ↩
- Google states there are no additional requirements or special optimizations to appear in AI Overviews or AI Mode; a page must be indexed and eligible to show with a snippet. Google Search Central, Dec 10, 2025; checked Sep 27, 2026. AI features and your website ↩
- OpenAI's help center says a website is eligible for ChatGPT search results only if it allows OAI-SearchBot and its host or CDN allows OpenAI's published search-bot IPs; placement is not guaranteed. OpenAI documentation, checked Sep 27, 2026. Searching the web with ChatGPT (OpenAI Help Center) ↩
- Google says no special schema.org markup or AI text files are needed for AI Overviews/AI Mode, but recommends structured data match visible text and that Business Profile and Merchant Center information be up to date. Google Search Central, Dec 10, 2025; checked Sep 27, 2026. AI features and your website ↩