Almost every company says it has backups. Fewer can say how much data they would lose if the database server failed right now, how long it would take to be running again, and when they last proved it by restoring. Those are the questions that matter during an incident. A backup that exists but has never been restored, sits in the same account as production, or takes two days to bring back does not protect the business the way people assume.
Start with RPO and RTO
| Term | Question it answers | Example |
|---|---|---|
| RPO (Recovery Point Objective) | How much data can we afford to lose? | Fifteen minutes for an order system, one day for an internal wiki |
| RTO (Recovery Time Objective) | How long can the application be down? | One hour for a customer portal, one working day for a reporting tool |
Agree on these numbers with the business owner of each application, because they drive the cost. A daily backup gives an RPO of up to 24 hours. Losing a full day of orders is rarely acceptable, so transaction systems usually need continuous database backups. An RTO of one hour means the restore must be scripted and practised, because nobody rebuilds a server from memory in sixty minutes during a crisis.
Follow the 3-2-1 Rule
- Keep at least three copies of important data: the production data and two backups.
- Store them on at least two different types of storage or services.
- Keep at least one copy off-site, in another region, another cloud account, or another provider.
- A common extension adds one immutable or offline copy that cannot be changed, and zero errors in restore tests.
Back Up What the Application Actually Needs
A working restore needs more than a database dump. List everything required to bring the application back: the databases, uploaded files in storage, configuration and environment variables, secrets such as API keys, DNS records, TLS certificates, and the infrastructure definition. Application code is usually safe in Git, but the build artifacts and container images used in production should also be reproducible. Infrastructure written as code, with Terraform or similar tools, turns rebuilding servers into a repeatable step instead of a long manual job.
Database Backups and Point-in-Time Recovery
Logical dumps such as pg_dump or mysqldump are simple and portable, but restoring a large database from a dump can take hours, and you lose everything since the last dump. Point-in-time recovery combines a periodic full backup with a continuous stream of transaction logs, such as WAL archiving in PostgreSQL or binary logs in MySQL, so you can restore to any moment, for example one minute before someone ran a bad delete. Managed databases such as Amazon RDS offer automated backups with point-in-time recovery, with a retention period you can set. Keep dumps as an extra, portable copy, and use point-in-time recovery for the RPO.
Protect Backups from Ransomware and Mistakes
- Store at least one copy in a separate account or project with different credentials, so an attacker who controls production cannot delete the backups.
- Use immutable storage such as S3 Object Lock or a backup vault lock, so backups cannot be deleted or changed during the retention period.
- Encrypt backups and keep the encryption keys where a restore can still access them during an incident.
- Limit who can delete backups or change retention, and alert when it happens.
- Keep enough history. Ransomware or data corruption is sometimes discovered weeks later, so a seven-day retention may not reach a clean copy.
Choose a Recovery Strategy
| Strategy | How it works | Typical RTO and cost |
|---|---|---|
| Backup and restore | Backups are copied to another region, and infrastructure is created only when needed | Hours, lowest cost |
| Pilot light | The database is replicated to another region, while application servers stay off until a disaster | Tens of minutes to an hour, low cost |
| Warm standby | A smaller copy of the full system runs in another region and is scaled up during a disaster | Minutes, moderate cost |
| Multi-site active-active | Two or more regions serve traffic at the same time | Near zero, highest cost and complexity |
Most business applications do well with backup and restore or pilot light, combined with infrastructure as code. Multi-site setups make sense for systems where every minute of downtime costs a lot, and they need a team that can operate that complexity.
Test Restores on a Schedule
- 1Write a recovery runbook with the exact steps, commands, and people involved, and store it outside the systems it describes.
- 2Restore the database to a separate environment every month and check that the application can read it.
- 3Measure how long the restore took and compare it with the RTO.
- 4Once or twice a year, run a full exercise: rebuild the application in another region from backups and infrastructure code.
- 5Update the runbook after every test with what was missing or slow.
Monitor the Backups Themselves
Backups fail quietly: a full disk, an expired credential, or a script that stopped after a server change. Send an alert when a backup job fails, and also when an expected backup does not arrive, using a heartbeat check. Watch backup size too. A backup that suddenly shrinks can mean the job is only saving part of the data.
Ask one question in your next operations meeting: when did we last restore production data, and how long did it take? If nobody knows, schedule the first restore test before adding any new backup tool.
Key takeaways
- Agree on RPO and RTO for each application, because those numbers decide the backup design and cost.
- Follow the 3-2-1 rule and keep at least one immutable copy in a separate account.
- Back up everything needed to rebuild the application, including files, configuration, secrets, and infrastructure code.
- Use point-in-time recovery for transactional databases, and keep dumps as an extra portable copy.
- Test restores on a schedule, measure the time against the RTO, and alert when backups fail or go missing.


