All terms

Search Tactic

Crawler policy

Also known as: AI crawler policy, robots.txt policy, bot access policy

A crawler policy is the set of rules a site publishes and enforces to decide which automated agents may access its content, how often, and for what purpose. It typically combines robots.txt directives, server side rules and terms of use, and increasingly distinguishes search crawlers from AI training and assistant crawlers. The policy defines both what is allowed and what happens when a bot ignores it.

What it is

It is a deliberate position on bot access rather than a default file left untouched. A policy names specific user agents, sets allow and disallow paths, and may apply rate limits or authentication to expensive endpoints. Many organisations now split decisions by purpose: indexing for search, retrieval for AI answers, and bulk collection for model training.

Why it matters

AI assistants and answer engines can only cite what they are permitted to fetch, so an overly broad block can quietly remove a brand from AI answers while a permissive setup may allow content to be reused with no traffic in return. Crawler decisions therefore shape visibility, brand mentions and referral volume at the same time. They also affect infrastructure cost, because aggressive bots consume real bandwidth and compute.

How it works

Practitioners audit server logs to see which agents actually visit, then write robots.txt rules per user agent and back them up with edge rules, verified bot allow lists and rate limits. Sensitive or low value paths such as internal search, faceted filters and checkout flows are disallowed, while key content and reference pages are kept open. The policy is documented, reviewed as new agents appear, and monitored for compliance.

When it applies

It applies to any public site, and needs explicit review whenever AI crawlers grow in your logs, a new content licensing position is taken, or crawl activity starts affecting performance.

Examples

  • A publisher allows search indexing crawlers but disallows a named AI training crawler while keeping an assistant retrieval agent allowed.
  • An ecommerce site blocks crawling of faceted filter URLs to protect crawl budget and server capacity.
  • A B2B site adds edge rate limiting after log analysis shows one unverified scraper generating a large share of requests.

How it is measured

  • Crawl requests and bytes served by user agent over time
  • Share of AI assistant answers citing the site, tracked before and after policy changes
  • Blocked or rate limited request volume and any resulting error rates
  • Coverage of key pages that remain crawlable in robots.txt testing

Related terms in Search Tactic

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

How ChatGPT Shortlists Software Brands downloads straight away. No spam, unsubscribe anytime.