Skip to content
ceeme

Which AI crawlers do I allow on my site?

This generator writes a robots.txt stating per AI crawler whether it may read your site. The list comes from the same register as our measurement — 23 crawlers, of which 15 can be verified against a published IP list or through DNS. The file is assembled entirely in your browser: no address travels to our server, because there is nothing to fetch.

By Njusja Orban, founder of ceeme3 min read

Measure your own company → Rather talk it through first

Pick a setting

Crawler by crawler

Advanced
robots.txt

      

In short

What the file does
It says per crawler who may come in. It is the only one of the three files on your domain a crawler is meant to obey.
Training versus live
Some crawlers collect for model training, others fetch during an answer. Blocking both is a different choice from blocking one.
The price of closing
What a crawler may not read ends up in no answer at all. Good copy does not undo that — nobody reads it.

Which setting suits me?

If you want to be found in AI answers: allow everything. That is the setting at the top, and the only one that costs nothing — a crawler reading your site uses bandwidth and nothing else.

The second setting is for anyone who objects to their texts training a model but still wants to appear in the answers. That is a defensible position and not a middle road: at some providers it is the same crawler, so in practice you still choose.

What is the difference between a training crawler and a live crawler?

A training crawler collects text that later ends up in a model; what it fetches today surfaces months later. A live crawler fetches during the conversation: what it finds is in the answer someone is reading now.

For short-term visibility the second kind counts, and it is also the one whose effect you can measure. The first is a long-term choice and a position on what you think of your own texts.

What does this file not do?

It governs access and nothing else. Whether you are read is one question; whether you are named is another, and it is not answered here.

What this tool doesWhat it does not do
Writes a valid robots.txt with a rule per crawlerDoes not place the file — that is you or your developer
Shows per crawler whether it can be verifiedDoes not check whether they actually visit — that needs your server log
Works entirely in your browserEnforces nothing: a crawler ignoring the file is not stopped by it

Frequently asked questions

Do I lose my Google ranking if I block AI crawlers?

Not automatically: the crawler maintaining the classic index is not the one collecting for AI, and the generator always lets the first one in. Do watch what you add by hand — one over-broad rule hits both, and that is the most common way a site accidentally drops out of the index.

Why are there more crawlers here than elsewhere?

Because the list comes from our own register and not from an article. There are 23 of them, each with whether it can be verified; for 8 of them that is not possible today, and it says so, because a name anyone can type is no proof of a visit.

Read on

Free, no account and no card.