A backup you’ve never restored is a hope
Every platform says it backs your data up. Very few can tell you the last time they put one back. We think the second sentence is the only one that counts, so when we built backups for AffinEQ we built the restore first, and we made it run every week.
What we built
Each night a service dumps the platform database and every project’s database. It encrypts each dump, records how many rows were in every table at the moment of the dump, and keeps a copy on the server and, once a bucket is configured, a second copy off it. Once a week it does the part that matters: it picks a project, restores its latest backup into a scratch database, checks the row counts match, and then reads the data back as the project’s own users. Not as an administrator, who can read anything, but through the same roles your app uses, which is where a bad restore tends to show up.
That last step is the point. A restore that works for the superuser and fails for the app is a backup that looks fine on the dashboard and fails on the worst day.
The first bug: the backup that protected itself
Backups pile up, so something has to delete the old ones, keeping the last 14 days, 8 weeks and 12 months. Our first version counted every run on the disk. That sounds right until you picture a push to the off-host copy that dies halfway, because a phone network, a power cut or a full disk interrupted it. Now there is a half-written run on disk, and it is the newest, so the pruning logic protected it as the latest backup while the last complete one, now older, aged out underneath.
A network that failed at the wrong moment, a few nights in a row, would have deleted every good backup we owned and left us holding fragments. We found it with a test that simulated exactly that push, not by looking. The fix is boring and strict: only runs that finished and were indexed count, the newest complete run is always kept, and an orphan is only swept once it is old enough that it cannot be a push still in progress.
The second bug: tests that cleaned up after the door was shut
Our integration tests create databases, roles and storage buckets, and delete them at the end. Nine of them never did. In Go, a defer runs before a t.Cleanup, and we had closed the database connection in a defer that ran first, so the cleanups that needed it ran against a closed connection and quietly did nothing.
The cost was invisible for weeks, because our CI starts from a fresh database each time. On the local machine we eventually counted 67 leftover project databases and 750 roles, and the object store ran out of room for new buckets. It would not have broken a customer. It did hide the fact that our tests did not tidy up, which is the kind of thing you only learn when something else, in this case the drill, makes you look.
The third bug: a database that cannot be restored
The drill failed on a project in our test rig, and the reason was subtle. A few rows referred to rows that did not exist, which a foreign key is supposed to make impossible. They were there because we had inserted them with checks switched off while testing. A dump contains them happily. A restore does not: it cannot add the foreign key back on top of data that breaks it, and with strict error handling it stops.
A real database can get into the same state in the same way, through a bulk load, a migration or a quick fix with constraints disabled, and it will sit there looking healthy until the day you need the backup. The drill found it in the test rig before it could find it in someone’s production data. That is its whole job.
Where it stands, and what it doesn’t do
It is running. The nightly backups happen by themselves, and every drill so far has passed, including a restore of a real project from backup in about four seconds, with its own roles reading the data afterwards.
We will be plain about the limits, because that is the point of publishing this:
- The off-host copy is built and tested but not switched on, because it needs a storage bucket we have not yet configured. Until then, the backups live on the same server they protect. That is better than nothing and not good enough, and our own tooling reports it as a failure.
- Backups are nightly, so the most you could lose is about a day. Point-in-time recovery, which narrows that to minutes, is planned and not built.
- You cannot yet press a button to download your data. It is plain Postgres, so
pg_dumpworks, but we want a better answer than that.
A backup you have never restored is a hope. A restore you run every week is a fact you can check.
If you want to see how we think about the rest of the platform, the proof section lists the commands we run to check our own claims.