Strategy & planning

Should I block AI crawlers in robots.txt?

Short answer
Not if you want to appear in AI answers. Blocking GPTBot, ClaudeBot, PerplexityBot, or Google-Extended removes your pages from the retrieval pool those engines draw on, which removes you from the answers. Blocking is a defensible choice for publishers protecting licensable archives, and a self-inflicted wound for almost everyone else.

Separate the two decisions

There are two distinct questions: may your content train future models, and may it be retrieved to answer questions now. Different user agents govern each, and conflating them is the most common mistake.

Retrieval bots such as OAI-SearchBot, ChatGPT-User and Perplexity-User fetch pages to answer a live question and cite the source. Blocking those is blocking referrals, not protecting assets.

A sane default

Allow the retrieval and search agents explicitly, keep private or per-user paths disallowed for everyone, and decide on training crawlers based on whether your content is a licensable asset in its own right.

Whatever you choose, check what is actually deployed. Blocking AI crawlers by accident — through a default rule, a CDN bot-protection setting, or a firewall preset — is far more common than blocking them on purpose.

See where you actually stand

We ask AI engines the questions your buyers ask and show you whether you were named, who was named instead, and which sources they cited. Free.

Related questions

More on strategy & planning

Last reviewed 2026-07-30.