Restore Drill
A backup that has never been restored is a hope, not a backup. Every six hours this agent restores the newest offsite backup of each HetOps database into a throwaway copy, proves it, and destroys it.
The problem
Every HetOps product runs on one server, and every night its database goes offsite to Cloudflare R2. A backup job that exits with success proves only that it ran. The file can still be empty, weeks old because the job quietly stopped, or missing the table that matters, and the usual place to find out is the day you need it.
How it works
- Findthe newest backup; fail if too old
- Restoretemp SQLite file or throwaway Postgres 17
- Proveintegrity, tables, rows, freshness
- Destroythen record the result and time
- Tellevidence page, health API, alerts

Decisions that matter
- Never near the live database. PostgreSQL dumps restore into a server the agent starts with
initdbin a temporary directory, listening on a Unix socket only, with no TCP port. It needs no Docker socket: a container that can reach it can control the host. - A read-only key scoped to one bucket. The agent can list and download; it cannot delete or overwrite a backup.
- No AWS SDK. The whole storage client, Signature V4 signing included, is about seventy hand-written lines, tested against the example signatures AWS publishes, so the whole agent has three runtime dependencies.
- Alerts on change, not on every pass: when a drill starts failing, when it recovers, and once a day while it stays broken. The last state survives a redeploy, so a deploy never re-alerts.
- Public proof, private detail. The page shows pass or fail, check names, restore times and backup age. Row counts, query results and storage URLs stay in the agent's logs.
What it caught
- Day one
The first DNS Intelligence drill failed, correctly. It expected a
watchestable that the real schema callsalerts. The backup was fine and the assumption was wrong, which is exactly the kind of assumption nobody checks until a restore. - A false alarm
After a redeploy, the status page reported the backups down for ten minutes: the last backup result lived only in memory. DNS Intelligence and Radar Cloud now keep it on disk next to the snapshots, so health is right from the first second after a deploy.
- Every six hours since
Three databases, two engines, restored and checked. The Umami PostgreSQL dump goes from download to a verified server in under two seconds.