All work

Restore Drill

A backup that has never been restored is a hope, not a backup. Every six hours this agent restores the newest offsite backup of each HetOps database into a throwaway copy, proves it, and destroys it.

Role
Author and operator
Stack
Node.js, SQLite, PostgreSQL 17, S3-compatible storage
Status
Open source, live at drill.hetops.dev, 0.3.0

The problem

Every HetOps product runs on one server, and every night its database goes offsite to Cloudflare R2. A backup job that exits with success proves only that it ran. The file can still be empty, weeks old because the job quietly stopped, or missing the table that matters, and the usual place to find out is the day you need it.

How it works

  1. Findthe newest backup; fail if too old
  2. Restoretemp SQLite file or throwaway Postgres 17
  3. Proveintegrity, tables, rows, freshness
  4. Destroythen record the result and time
  5. Tellevidence page, health API, alerts
The live evidence page: 3 of 3 backups restored, a verified stamp, and the four steps of the last drill.
The evidence page at drill.hetops.dev. The live section on the home page reads the same API.

Decisions that matter

What it caught

  1. Day one

    The first DNS Intelligence drill failed, correctly. It expected a watches table that the real schema calls alerts. The backup was fine and the assumption was wrong, which is exactly the kind of assumption nobody checks until a restore.

  2. A false alarm

    After a redeploy, the status page reported the backups down for ten minutes: the last backup result lived only in memory. DNS Intelligence and Radar Cloud now keep it on disk next to the snapshots, so health is right from the first second after a deploy.

  3. Every six hours since

    Three databases, two engines, restored and checked. The Umami PostgreSQL dump goes from download to a verified server in under two seconds.

Got a production problem worth solving?