plugins

BareProxy AI Crawler Control Plugin: Allow, Block, Rate-Limit or Charge Each AI Bot

Status: planned, number 10 of 21 in BareProxy’s build order. The plugins are built easiest first, and this one is a day or two of coding: bot rules, address lists fetched and refreshed on a timer, counters. It comes after the Link previews plugin. This post describes what it will do, and it will be updated as it is built.

AI companies crawl the web to train models and to answer questions with fresh pages. GPTBot, ClaudeBot, PerplexityBot, CCBot, Bytespider and a growing list of others now make up a real share of the requests many sites serve. A site owner’s choices so far are thin. robots.txt is a polite request that some bots honor and some don’t, and it can’t say “yes, but slower” or “yes, if you pay”. Blocking by user agent is easy to get around, because anyone can send any user agent.

AI crawler control gives each bot its own answer, at the proxy, before the request reaches the site.

What It Will Do

  • Allow, block or slow down each bot, by name. A rule can let GPTBot in, keep Bytespider out, and hold PerplexityBot to a few requests a minute.
  • Check that a bot is who it says it is. Where an operator publishes the address ranges its crawler uses, the plugin checks the request’s address against them, and keeps the list fresh on a timer. A request that claims to be GPTBot from an address outside OpenAI’s ranges is treated as an unknown bot.
  • Send it to the payment gate. A rule can hand a bot to the AI crawler payment gate plugin, which answers 402 Payment Required with the terms and lets a crawler in once it pays.
  • Keep a log of who is scraping what. Every bot request is counted by bot, path and answer, so the question “how much of my traffic is AI crawlers, and what are they reading?” has a plain answer.

A config will look something like this:

plugin crawlers /etc/bareproxy/plugins/ai-crawler-control.wasm
  config /etc/bareproxy/plugins/crawlers.json
  allow-http openai.com:443
  on-error open

site example.com
  use crawlers
  route /* -> files /var/www/example/public

on-error open matters here: if the plugin ever fails, the site keeps serving people, and only the bot rules pause.

Why It Matters Now

It answers a question nearly every site owner is asking right now. It needs the request’s headers and address, a timer, shared counters and the request record, and nothing heavy, which is why it sits in the first half of the build order.

Every decision it makes lands in the request’s record. why will say, for any request: this was ClaudeBot, its address checked out, the rule allowed it, and it was the 41st request from it this hour.

Part of a Pair

Crawler control works together with the RenderCache plugin, which comes later in the build order because it needs a headless browser beside BareProxy. Crawler control decides whether a bot gets in and on what terms. RenderCache decides what an admitted bot sees: a fully rendered page instead of an empty JavaScript shell.

The whole program is on the plugins page.