How to Run a Full Backup Restore Test
A green backup job means files were written somewhere. It does not mean you can rebuild the service from them, or how long that takes. The only proof is a restore. A full restore test starts from nothing but the backups and the documentation, rebuilds the service in an isolated environment, checks that it works, and measures how long it took.
Define scope and targets
List every piece of state the service needs:
- databases,
- object storage buckets,
- Kubernetes resources and persistent volumes,
- secrets and encryption keys, including the keys that decrypt the backups,
- configuration that lives outside git: DNS, SaaS settings, feature flags.
For each one, write down the recovery point objective (how much data you can lose) and recovery time objective (how long you can be down). The test measures the real values against them.
Prepare an isolated environment
Restore into a separate cloud account, project or cluster. Do not restore next to production.
- No network path to production databases and queues.
- Outbound side effects switched off: email, payments, webhooks, scheduled jobs that call third parties.
- Restore with the credentials that would be used in a real disaster, not with an admin account. A restore role missing permissions is a common finding.
- Have someone other than the backup's author follow the written procedure. That is how you find the steps that only exist in one person's head.
Record the start time and a timestamp for every step.
Restore PostgreSQL with pgBackRest
Check what is available, then restore to a point in time into a fresh data directory:
sudo -u postgres pgbackrest --stanza=main info
install -d -m 700 -o postgres /var/lib/postgresql/16/restore-test
sudo -u postgres pgbackrest --stanza=main restore \
--pg1-path=/var/lib/postgresql/16/restore-test \
--type=time --target="2026-10-04 23:00:00+00" \
--target-action=promote \
--archive-mode=off
sudo -u postgres /usr/lib/postgresql/16/bin/pg_ctl \
-D /var/lib/postgresql/16/restore-test -o "-p 5433" -l /tmp/restore-test.log start
--archive-mode=off matters: without it, a promoted test cluster that still has the production archive_command can push WAL into the production repository. --target-action=promote opens the database once the target time is reached. The default is pause, which leaves it read-only and waiting.
If the repository is encrypted, the test also proves whether repo1-cipher-pass is stored somewhere you can reach when the production host is gone. Between full tests, pgbackrest --stanza=main verify checks that the backups and the WAL archive in the repository are valid.
For RDS, the equivalent is a point-in-time restore into a new instance. Pass the subnet group and security groups explicitly, otherwise the instance gets defaults:
aws rds restore-db-instance-to-point-in-time \
--source-db-instance-identifier my-app-prod \
--target-db-instance-identifier my-app-restore-test \
--restore-time 2026-10-04T23:00:00Z \
--db-subnet-group-name restore-test \
--vpc-security-group-ids sg-0123456789abcdef0
Restore Kubernetes with Velero
Point Velero in the test cluster at the same backup storage location, read-only, so the test cannot delete or overwrite production backups:
velero backup-location create prod-backups --provider aws \
--bucket my-velero-backups --config region=eu-central-1 \
--access-mode ReadOnly
velero backup get
velero backup describe daily-20261004 --details
velero restore create my-app-restore-test \
--from-backup daily-20261004 \
--include-namespaces my-app \
--namespace-mappings my-app:my-app-restore-test
velero restore describe my-app-restore-test --details
velero restore logs my-app-restore-test
Read the backup description before restoring. Velero only has volume data if volume snapshots or file system backup were configured. Without them you get the manifests and empty volumes. Cloud volume snapshots are also tied to a region and usually to an account, so a restore into another region needs snapshots copied there first.
Verify, not just start
"The pods are running" is not a result. Check:
- Data freshness. The newest record shows your real recovery point:
SELECT max(created_at) FROM orders; - Data completeness. Row counts of key tables compared with production at the target time, and every expected database and schema present.
- Schema version. The migration table (
databasechangelog,schema_migrationsor similar) matches the version of the application you deploy. - Application behavior. A smoke test through the real API: log in, read, write, run one background job.
- Files. Checksums or object counts for restored buckets and volumes.
Measure and write down the gaps
Actual recovery time is the time from the start of the test until the smoke test passes, including time spent finding credentials, waiting for access and reading documentation. Compare it with the target.
Typical gaps to look for:
- tables, databases or buckets excluded from backup,
- backup encryption keys or passphrases stored only on production hosts,
- restore roles without the permissions they need,
- snapshots in the same account or region as the thing they protect,
- configuration made by hand and never captured in code,
- restored services still pointing at production hostnames,
- secrets in an external secret store that is not backed up.
Each gap becomes a ticket with an owner, and the runbook is updated before the test is closed.
Make it routine
Run a full restore test on a fixed schedule and after major infrastructure changes. Between full tests, automate the cheap part: a scheduled job that restores the latest database backup into a scratch instance, runs a few queries, reports the newest record's timestamp, and alerts if anything fails.
Checklist
- Inventory of all state, with RPO and RTO for each.
- Isolated environment, no path to production, side effects off.
- Restore with disaster-time credentials, following the written runbook.
- pgBackRest restores with
--archive-mode=off; Velero reads backups through a read-only location. - Verify freshness, completeness, schema version and application behavior.
- Measure real recovery time and turn every gap into a ticket.
- Automated daily restore checks between full tests.
