<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Bots on BareProxy.com</title>
    <link>https://bareproxy.com/tags/bots/</link>
    <description>Recent content in Bots on BareProxy.com</description>
    <generator>Hugo</generator>
    <language>en-us</language>
    <lastBuildDate>Thu, 08 Oct 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://bareproxy.com/tags/bots/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>BareProxy AI Crawler Control Plugin: Allow, Block, Rate-Limit or Charge Each AI Bot</title>
      <link>https://bareproxy.com/bareproxy-ai-crawler-control-plugin-allow-block-rate-limit-or-charge-each-ai-bot/</link>
      <pubDate>Thu, 08 Oct 2026 00:00:00 +0000</pubDate>
      <guid>https://bareproxy.com/bareproxy-ai-crawler-control-plugin-allow-block-rate-limit-or-charge-each-ai-bot/</guid>
      <description>&lt;p&gt;&lt;strong&gt;Status: planned, number 10 of 21 in BareProxy&amp;rsquo;s build order.&lt;/strong&gt; The plugins are built easiest first, and this one is a day or two of coding: bot rules, address lists fetched and refreshed on a timer, counters. It comes after the Link previews plugin. This post describes what it will do, and it will be updated as it is built.&lt;/p&gt;&#xA;&lt;p&gt;AI companies crawl the web to train models and to answer questions with fresh pages. GPTBot, ClaudeBot, PerplexityBot, CCBot, Bytespider and a growing list of others now make up a real share of the requests many sites serve. A site owner&amp;rsquo;s choices so far are thin. robots.txt is a polite request that some bots honor and some don&amp;rsquo;t, and it can&amp;rsquo;t say &amp;ldquo;yes, but slower&amp;rdquo; or &amp;ldquo;yes, if you pay&amp;rdquo;. Blocking by user agent is easy to get around, because anyone can send any user agent.&lt;/p&gt;</description>
    </item>
    <item>
      <title>BareProxy AI Crawler Payment Gate Plugin: 402 Payment Required for Bots That Want the Content</title>
      <link>https://bareproxy.com/bareproxy-ai-crawler-payment-gate-plugin-402-payment-required-for-bots-that-want-the-content/</link>
      <pubDate>Thu, 08 Oct 2026 00:00:00 +0000</pubDate>
      <guid>https://bareproxy.com/bareproxy-ai-crawler-payment-gate-plugin-402-payment-required-for-bots-that-want-the-content/</guid>
      <description>&lt;p&gt;&lt;strong&gt;Status: planned, number 19 of 21 in BareProxy&amp;rsquo;s build order.&lt;/strong&gt; The plugins are built easiest first, and this one is about four or five days of coding: prices, proof of payment and books, on payment schemes still settling. It comes after the Response cache plugin. This post describes what it will do, and it will be updated as it is built.&lt;/p&gt;&#xA;&lt;p&gt;AI companies need fresh, well-written pages, and a growing number of publishers want to be paid for them. The web has had a status code for this since 1997 that almost nobody used: 402 Payment Required. It is finally getting a job. A site answers a crawler with 402 and its terms; the crawler pays and comes back with proof; the site serves the page.&lt;/p&gt;</description>
    </item>
    <item>
      <title>BareProxy Bot Challenges Plugin: A Proof-of-Work Check Before Scrapers Reach the Site</title>
      <link>https://bareproxy.com/bareproxy-bot-challenges-plugin-a-proof-of-work-check-before-scrapers-reach-the-site/</link>
      <pubDate>Thu, 08 Oct 2026 00:00:00 +0000</pubDate>
      <guid>https://bareproxy.com/bareproxy-bot-challenges-plugin-a-proof-of-work-check-before-scrapers-reach-the-site/</guid>
      <description>&lt;p&gt;&lt;strong&gt;Status: planned, number 11 of 21 in BareProxy&amp;rsquo;s build order.&lt;/strong&gt; The plugins are built easiest first, and this one is a day or two of coding: a challenge page, a check and a signed cookie. It comes after the AI crawler control plugin. This post describes what it will do, and it will be updated as it is built.&lt;/p&gt;&#xA;&lt;p&gt;Not every bot announces itself. Scrapers that ignore robots.txt usually send a browser&amp;rsquo;s user agent and spread their requests over thousands of addresses, so neither user-agent rules nor rate limits catch them. Small sites, code forges and wikis have been hit hard by this, sometimes to the point of going offline.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
