Some folks want their archives available so they can link to them, but don't want them to come up in Google results.
Some folks want to prohibit LLMs and other scrapers from crawling their archives.
Both of these are eminently reasonable! We tried to handle a lot of this with a very engineer-brained approach: give you a big ol' textarea where you can type user agent strings to block.
People do not know or care about user agent strings, we realized. So we replaced all of it with a Crawling add-on in Settings → Archives, holding one toggle per kind of crawler.
We looked at usage (and our support volume) and found the two biggest reasons/avenues that y'all care about, plus a catch-all for everyone else:
| Toggle | Examples |
|---|---|
| Search engines | Googlebot, Bingbot, DuckDuckBot, Applebot |
| AI crawlers | GPTBot, ClaudeBot, CCBot, Google-Extended, PerplexityBot, and a few dozen more |
| Everything else | Scrapers, SEO tools like AhrefsBot, archivers, and bots that don't say who they are |
The docs have the full rundown of what each toggle does under the hood.
