Features
A Backup You Have Not Restored Is a Guess
What a real restore test involves, and the specific moment a backup turned out to hold one row instead of forty-two.
23 September 2026

The Restore Test That Reveals Gaps
When you test database backup restore procedures, you are verifying that the data you think you have actually exists and functions. A backup file that exists on disk is not a backup; it is a guess until you have proven it can be turned back into a live, queryable state. The difference between a file and a safety net is the act of restoration itself.
Many developers assume that if the backup job runs without error, the data is safe. This is a dangerous fallacy. A process can successfully write a zero-byte file, a corrupted archive, or a snapshot that missed the last transaction. The only way to know is to stop pretending the backup is safe and actually bring it up in an isolated environment. If you cannot restore it, you do not have a backup. You have a hope.
Why One Row Is Not Enough
The specific failure mode that haunts small operations is the partial write. Imagine a scheduled task that dumps a table with forty-two rows. The process starts, writes the header, and then encounters a temporary lock or a disk buffer issue. It might exit with a success code, or it might write a file that contains only the first row. Without a verification step, you will not know the difference. The file exists. The timestamp is recent. The size looks plausible if you are not checking the metadata.
In a setup where a single PostgreSQL 16 database serves multiple applications, the stakes are higher. If the dump is truncated, the restore will fail, or worse, it will succeed with incomplete data. The application might start, but it will be missing critical records. For a service like Instalaz, which relies on precise scheduling data, a missing row is not a minor inconvenience; it is a broken promise to the user. The restore test must check not just that the command finishes, but that the row counts match the source exactly.
Running the Test on a Single VPS
On a single VPS with 12GB of RAM and a 193GB disk, resources are finite. Running a restore test on the production machine is risky because processes on one box compete for memory and CPU. A large restore operation can cause the primary database to slow down or even crash if the disk I/O spikes. The standard practice is to restore the backup to a separate, isolated instance.
This requires a temporary PostgreSQL instance, often on a different port, that is not connected to the main application. You restore the backup file into this instance, run a series of checksums or simple count queries, and then shut it down. This proves the file is valid. If you skip this step, you are relying on the assumption that the backup tool is perfect. Tools fail. Disks fail. Network connections to remote storage fail. The isolated restore is the only proof that works.
Automating the Verification
Manual testing is too easy to forget. The verification must be part of the automated pipeline. Since scheduled work runs as systemd oneshot services with timers, the backup script should not just dump the database. It should also trigger a restore test in a sandbox environment.
The script should compare the row counts of the source tables against the restored tables. If there is a mismatch, the job should fail loudly. This alert should be treated with the same severity as a production outage. A backup that fails its restore test is not a backup. It is a liability. By automating this check, you ensure that every night, the system proves it can save you. This applies to any data-driven product, whether it is the scheduling data for Instalaz or the user profiles for a film publication. The data must be recoverable, or it does not matter how well it performs while it is alive.


