plugins

BareProxy RenderCache Plugin: Prerendered Pages for Search Bots, AI Crawlers and Link Previews

Status: planned, number 21 of 21 in BareProxy’s build order. The plugins are built easiest first, and this one is about one to two weeks of coding: the plugin plus a headless browser beside BareProxy, and pages that must render right. It comes after the Image optimizer plugin. This post describes what it will do, and it will be updated as it is built.

A JavaScript app shows a browser a nearly empty HTML page and builds the rest in the browser. People never notice. Crawlers do. Google can render JavaScript, slowly and on its own schedule. Most AI crawlers don’t run it at all: they fetch the HTML, find a <div id="root"> and a bundle of scripts, and move on. Link previews in chat apps and social networks do the same. For them, a single-page app is a blank page.

RenderCache fixes that at the proxy. When a search bot, an AI crawler or a link-preview fetcher asks for a page, BareProxy hands it a fully rendered HTML snapshot from a cache. Everyone else gets the live app, untouched.

How It Will Work

RenderCache comes in two parts, because rendering a page means running a real browser, and a browser doesn’t belong inside a sandboxed plugin.

  • The plugin, inside BareProxy, decides. It looks at each request before routing, recognizes bots and preview fetchers, and when one arrives it serves the snapshot from its own store on disk. A miss, or a snapshot past its age, goes to the renderer.
  • The renderer is a small separate program next to BareProxy that drives headless Chromium. The plugin calls it over HTTP, at an address its config line allows. It loads the page the way a browser would, waits until the page settles, and hands back the HTML.

Snapshots refresh on a schedule, or at once through a purge call when the content changes. A config will look something like this:

plugin rendercache /etc/bareproxy/plugins/rendercache.wasm
  config /etc/bareproxy/plugins/rendercache.json
  store 2GB
  allow-http 127.0.0.1:9222

site example.com
  use crawlers rendercache
  route /* -> app

why will show, for every bot request, whether it got a snapshot, how old it was, or why it went to the app instead.

One Rule It Won’t Break

A snapshot has to show what a person would see. Serving bots different content is cloaking, and search engines punish it. RenderCache renders the same page a visitor gets; it changes when the bot sees it, never what. Google calls this kind of dynamic rendering a workaround rather than a long-term answer for its own crawler, and that’s fair for Googlebot. AI crawlers and preview fetchers are the bigger reason to build it. They don’t render, and they aren’t about to start.

Part of a Pair

RenderCache works together with the AI crawler control plugin, which is built well before it, and the two tell one story: BareProxy decides what bots see, and on what terms. Crawler control decides whether a bot gets in. RenderCache decides what it gets.

The whole program is on the plugins page.