Guides / How the engines find and choose businesses
How ChatGPT, Gemini, Perplexity, and Claude find local businesses through search indexes, crawlers, and citations
The mechanics behind an AI recommendation: training data versus live search, each engine's documented crawler, what a citation means, and the checks a local business can run on its own site.
By Ben Hinton
, published
The question: How do ChatGPT, Gemini, Perplexity, and Claude find local businesses?
The short answer
When someone asks an AI assistant for a roofer or an air conditioning company, an engine with search looks it up: it searches the web, reads a handful of pages, and writes an answer from what it read, sometimes with links to those pages. Each engine reaches the web through crawlers and search indexes that it documents for site owners. What an engine can say about your business is limited to what it can reach and what the pages it reads say about you.
So there are three questions to ask about any local business. Can the engine's crawler reach the website? Is the business described on the other pages the engine reads, such as directories and review sites? And do those pages agree with each other?
Two ways an engine knows about you
A language model learns from a large body of text before it is released. That is its training data. OpenAI, for example, crawls content that may be used to train its models with a separate crawler called GPTBot1. Anthropic documents ClaudeBot2 as the bot that collects web content that could contribute to training its models. Training gives a model general knowledge: what a heat pump is, what a roof inspection involves, which cities are in the East Valley.
Training is a poor source for local recommendations. It is frozen at a point in time, and a contractor's phone number, reviews, and service area change. When a question needs current, local facts, the engines with search turn to live retrieval: they search an index of the web, fetch the pages that match, and ground their answer in those pages. That is the path that matters for being recommended, and it is the one each engine documents.
ChatGPT
OpenAI states that OAI-SearchBot3 is the crawler used to surface websites in ChatGPT's search features, and that sites that opt out of it are not shown in ChatGPT search answers, though they can still appear as navigational links. Its help center gives the eligibility rule in one line: allow OAI-Searchbot to crawl the site4, and confirm that the site's host or content delivery network allows traffic from OpenAI's published search-bot IP addresses.
OpenAI also works with outside search providers. Its help center for business workspaces says ChatGPT may share disassociated search queries with Bing5 to return web results. What ChatGPT reads can depend on another company's index as well as OpenAI's own crawling.
The training crawler and the search crawler are controlled separately. A site can decline GPTBot and still allow OAI-SearchBot, so opting out of training does not have to mean opting out of ChatGPT search.
Perplexity
Perplexity documents PerplexityBot6 as the crawler designed to surface and link websites in its search results, and says it is not used to crawl content for training AI foundation models.
Claude
Anthropic runs separate bots for training, for fetches a user asks for, and for search. Its search crawler is Claude-SearchBot7. Anthropic says that disabling it prevents Claude from indexing your content for search, which may reduce your site's visibility and accuracy in Claude's search results. As with OpenAI, a site can block the training bot and allow the search bot.
Google: AI Overviews, AI Mode, and Gemini
Google's documentation for site owners covers its AI features in Search, AI Overviews and AI Mode, and it is short. There are no additional requirements8 to appear in them, nor other special optimizations: a page has to be indexed and eligible to appear in Google Search with a snippet. Google also documents that both features may use query fan-out9, running several related searches across subtopics and data sources to build one response. The same page recommends keeping Business Profile information up to date, and says there is no special schema.org structured data10 you need to add. For a local business, Google's index and its Business Profile are the ground floor. We measure Gemini's answers directly, and its citation rate is in the list below.
What a citation means
A citation is a link an engine shows beside its answer, pointing at a page it used. It is the only direct evidence, from the outside, of what an engine read. It is also not always there.
Our latest weekly Phoenix measurement covers window and door questions asked of Claude and Gemini. In it, the share of answers that cited at least one website was:
ChatGPT and Perplexity will be added to the measurement, and their rates will join this list when they are measured.
An answer with no citation can still name businesses, and it still came from somewhere. The absence of a link just means you cannot see the source. That is why a serious measurement records both which businesses an answer names and which sites it cites, and treats them as separate things. Our window and door index alone is built from 8613 answers in its newest weekly run. We have not measured roofing or heating and air conditioning yet.
What the engines read about a local business
The retrieved pages fall into a few groups: the business's own website, directory and review profiles, Google's own business data, and articles that mention the business. Outside studies suggest directories carry real weight. BrightLocal found Yelp used as a source in 33%14 of its local test searches across ChatGPT Search, Gemini, Google AI Mode, and Perplexity. The sites that carry the most weight for Phoenix contractors specifically are in which websites AI assistants cite.
Reviews appear to count. SOCi reports that locations ChatGPT recommends average 4.3-star ratings15. That is a vendor's correlation across multi-location brands, not proof of a rule. It is still a reason to treat your review profiles as part of what the engines read about you.
What this means for an owner
Four things decide whether an engine can find you, and each can be checked:
- Crawler access. Your robots.txt and your host or CDN must let the search crawlers in: OAI-SearchBot, Claude-SearchBot, PerplexityBot, and Googlebot. Security settings that block unfamiliar bots can shut them out without anyone noticing. Our guide to llms.txt and AI crawler access covers the details.
- Index presence. Your important pages must be indexed by Google, and it is worth confirming Bing sees them too, given that OpenAI names Bing as a search provider.
- Consistent facts. Your name, phone, address or service area, hours, and lines of business should read the same on your site, your Business Profile, and every directory the engines cite. See structured data for AI visibility.
- Pages worth citing. A page that answers one homeowner question plainly, with specifics from your own work, gives an engine something to quote.
None of this guarantees a recommendation. It makes you findable, and then the only honest test is measurement: the same questions, asked of every engine, week after week.
Notes
- OpenAI's GPTBot crawls content for training its models and is controlled independently of OAI-SearchBot, so a site can block training while still appearing in ChatGPT search. OpenAI documentation, checked Sep 27, 2026. Overview of OpenAI Crawlers ↩
- Anthropic's ClaudeBot collects web content that could contribute to training its generative AI models; restricting it signals that a site's future materials should be excluded from Anthropic's training datasets. Anthropic, Apr 7, 2026; checked Sep 27, 2026. Does Anthropic crawl data from the web, and how can site owners block the crawler? ↩
- OpenAI's OAI-SearchBot crawler is what surfaces websites in ChatGPT search; sites that block it are not shown in ChatGPT search answers. OpenAI documentation, checked Sep 27, 2026. Overview of OpenAI Crawlers ↩
- OpenAI's help center says a website is eligible for ChatGPT search results only if it allows OAI-SearchBot and its host or CDN allows OpenAI's published search-bot IPs; placement is not guaranteed. OpenAI documentation, checked Sep 27, 2026. Searching the web with ChatGPT (OpenAI Help Center) ↩
- OpenAI documents that ChatGPT search sends disassociated search queries to Bing to return web results (stated explicitly for Enterprise and Edu workspaces). OpenAI documentation, checked Sep 27, 2026. ChatGPT search for Enterprise and Edu (OpenAI Help Center) ↩
- PerplexityBot is the crawler that surfaces and links websites in Perplexity's search results and is not used for model training; Perplexity recommends allowing it to appear in results. Perplexity documentation, checked Sep 27, 2026. Perplexity Crawlers ↩
- Anthropic runs three bots - ClaudeBot (training), Claude-User (user-directed fetches) and Claude-SearchBot (search indexing); blocking Claude-SearchBot may reduce a site's visibility in Claude's search results. Anthropic, Apr 7, 2026; checked Sep 27, 2026. Does Anthropic crawl data from the web, and how can site owners block the crawler? ↩
- Google states there are no additional requirements or special optimizations to appear in AI Overviews or AI Mode; a page must be indexed and eligible to show with a snippet. Google Search Central, Dec 10, 2025; checked Sep 27, 2026. AI features and your website ↩
- Google documents that AI Overviews and AI Mode may use 'query fan-out', issuing multiple related searches across subtopics and data sources to build one answer. Google Search Central, Dec 10, 2025; checked Sep 27, 2026. AI features and your website ↩
- Google says no special schema.org markup or AI text files are needed for AI Overviews/AI Mode, but recommends structured data match visible text and that Business Profile and Merchant Center information be up to date. Google Search Central, Dec 10, 2025; checked Sep 27, 2026. AI features and your website ↩
- Share of Claude's home-services answers that cited at least one website: 74%. Across 1 of 3 trades, each trade's last 4 weekly runs (1 run in all), Sep 21, 2026 to Sep 27, 2026: 86 answers. See the data ↩
- Share of Gemini's home-services answers that cited at least one website: 71%. Across 1 of 3 trades, each trade's last 4 weekly runs (1 run in all), Sep 21, 2026 to Sep 27, 2026: 86 answers. See the data ↩
- AI answers measured for window and door questions in Phoenix: 86. 86 answers from 2 engines, week of Sep 21, 2026. See the data ↩
- In BrightLocal's test of local searches across ChatGPT Search, Gemini, Google AI Mode and Perplexity, Yelp was used as a source in 33% of searches, and every platform used directories. BrightLocal, Jul 22, 2025; checked Sep 27, 2026. AI Search Makes Local Listings More Important Than Ever ↩
- Locations that ChatGPT recommends average a 4.3-star rating. SOCi, Jan 28, 2026; checked Sep 27, 2026. In AI-Driven Discovery, Few Brands Are Chosen, Most Disappear ↩