Skip to content

Free tool

Check which AI crawlers your robots.txt allows

Reads your robots.txt the way each answer engine reads it, and separates the two questions everyone conflates: whether you can be cited, and whether you can be used as training data.

Fetches one file — your robots.txt — and reads it the way each agent does.

Retrieval and training are not the same decision

A retrieval crawler fetches your page when somebody asks a question, and the answer cites you. A training crawler collects corpus. Blocking the second is a licensing choice more publishers make every year. Blocking the first removes you from answers people are reading today — usually by accident, because both were disallowed by a template.

The agents this tool checks, and what each is for
AgentWhat blocking it costs
GooglebotgooglebotGoogle Search and AI Overviews
OAI-SearchBotoai-searchbotChatGPT search results
ChatGPT-Userchatgpt-userChatGPT following a link to you in a live conversation
PerplexityBotperplexitybotPerplexity
ClaudeBotclaudebotClaude
GPTBotgptbotOpenAI training corpora
Google-Extendedgoogle-extendedGemini grounding
Applebot-Extendedapplebot-extendedApple Intelligence
CCBotccbotCommon Crawl, which most models train on

robots.txt is a claim. The audit measures the outcome.

The full audit requests your page as each agent and records the status code it actually receives, because bot management at the edge can allow a crawler in robots.txt and answer it with a 403 anyway. It also checks whether your content exists without JavaScript, which is what most AI crawlers see.

Run the full audit

AI visibility checks