Command Palette

Search for a command to run...

Hectal
PHASE 4Intermediate ~15 min· topic 5 of 5

Topic 4.5

Backups & Data Protection: AWS Backup, Object Lock & Restore Drills

In one line

Replication protects against hardware failure, not against mistakes, bad deploys, or ransomware, because those get replicated too. Real protection is backups: AWS Backup plans that cover EBS, RDS/Aurora, DynamoDB, EFS, and S3 with schedules and retention; copies to another account and region; immutability with Backup Vault Lock and S3 Object Lock; and regular restore drills that prove you can meet your RPO and RTO.

0/5 · 0%

Think of it like this

A photocopy of your house papers kept in the same cupboard (replication) burns in the same fire. A copy in a bank locker in another city (cross-region, cross-account backup) that even you can't shred early (vault lock) survives. And you only know it works if you once tried using it (restore drill).

Words you'll meet

New words in this topic, in plain English. Come back here whenever one feels fuzzy.

RPO
Recovery Point Objective: the maximum acceptable data loss, measured in time.
RTO
Recovery Time Objective: the maximum acceptable time to restore service.
Backup plan
An AWS Backup rule set defining schedule, retention, and copies.
Backup vault
An encrypted container in AWS Backup where recovery points are stored.
Vault Lock
A setting that makes backups in a vault impossible to delete before retention ends.
Object Lock
An S3 feature that prevents objects from being deleted or overwritten for a set period.

Step by step

01A backup plan by tag

Tiffin agrees an RPO of 1 hour for orders (Aurora PITR covers that) and 24 hours for everything else. A daily AWS Backup plan keeps 35 days of backups and copies each one to a locked vault in the backup account in Hyderabad (ap-south-2).

backup-plan.jsonwhole filejson

21:00 UTC is 02:30 IST, after the dinner rush. 555566667777 is the separate backup account.

{
  "BackupPlanName": "tiffin-daily",
  "Rules": [{
    "RuleName": "daily-0230-ist",
    "TargetBackupVaultName": "tiffin-prod",
    "ScheduleExpression": "cron(0 21 * * ? *)",
    "StartWindowMinutes": 60,
    "Lifecycle": { "DeleteAfterDays": 35 },
    "CopyActions": [{
      "DestinationBackupVaultArn": "arn:aws:backup:ap-south-2:555566667777:backup-vault:tiffin-locked",
      "Lifecycle": { "DeleteAfterDays": 90 }
    }]
  }]
}
terminal
$ PLAN=$(aws backup create-backup-plan --backup-plan file://backup-plan.json --query BackupPlanId --output text)
aws backup create-backup-selection --backup-plan-id $PLAN --backup-selection '{"SelectionName":"tagged","IamRoleArn":"arn:aws:iam::111122223333:role/aws-backup","ListOfTags":[{"ConditionType":"STRINGEQUALS","ConditionKey":"backup","ConditionValue":"daily"}]}' --query SelectionId --output text
── expected output ──
a1b2c3d4-5e6f-7a8b-9c0d-1e2f3a4b5c6d
A backup plan by tagdiagram
Rendering diagram…

02Locking the backup vault

In the backup account, the vault is locked in compliance mode with a minimum retention. After the 3-day grace period, nobody can remove the lock or delete recovery points early, not even the account's root user.

terminal
$ aws backup put-backup-vault-lock-configuration --backup-vault-name tiffin-locked --min-retention-days 30 --max-retention-days 365 --changeable-for-days 3 --region ap-south-2
aws backup describe-backup-vault --backup-vault-name tiffin-locked --region ap-south-2 --query '[Locked,MinRetentionDays,LockDate]' --output text
── expected output ──
True 30 2026-10-05T09:40:12.000Z
Before LockDate you can still change or remove the lock. Test in a sandbox first, because after that it's permanent.

03The restore drill

On the first Monday of each month Priya restores last night's Aurora backup into an isolated subnet, runs data checks, and records how long it took. This month: 41 minutes against an RTO of 2 hours.

terminal
$ aws backup start-restore-job --recovery-point-arn arn:aws:backup:ap-south-1:111122223333:recovery-point:6f1e... --iam-role-arn arn:aws:iam::111122223333:role/aws-backup \
--metadata '{"DBClusterIdentifier":"drill-orders-2026-10","Engine":"aurora-postgresql","DBSubnetGroupName":"drill-isolated"}' --query RestoreJobId --output text
aws backup describe-restore-job --restore-job-id 3c4d5e6f-... --query '[Status,CompletionDate]' --output text
psql -h drill-orders-2026-10.cluster-c1a2b3c4d5e6.ap-south-1.rds.amazonaws.com -U tiffin_admin -d tiffin -c "select count(*), max(created_at) from orders"
── expected output ──
3c4d5e6f-7a8b-9c0d-1e2f-3a4b5c6d7e8f
COMPLETED 2026-10-05T10:21:33Z
count | max
---------+------------------------
1842213 | 2026-10-04 20:58:41+00
The newest order is from just before the backup, as expected. Then delete the drill cluster.

Break it on purpose

Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.

Break #1

Replication isn't backup

A migration script runs DELETE FROM menu_items WHERE restaurant_id = restaurant_id (a typo, always true) on the writer. The team relies on Multi-AZ and a read replica.

terminal
$ psql -h tiffin-orders.cluster-ro-c1a2b3c4d5e6.ap-south-1.rds.amazonaws.com -U app -d tiffin -c 'select count(*) from menu_items'
── what you'll see ──
count
-------
0
# the reader, the standby, and every replica faithfully copied the delete within milliseconds

Myth vs fact

Myth

S3 has eleven nines of durability, so I don't need backups of it.

Fact

Durability protects against AWS losing data, not you deleting or overwriting it, or an attacker encrypting it. Use versioning, replication to another account, and Object Lock for critical data.

Pro corner

Extra depth for experienced readers. New to this? Skip it for now and come back later.

  • ▸

    Treat your backup account like a vault: separate credentials, no CI access, SCPs that deny backup:DeleteRecoveryPoint and backup:PutBackupVaultAccessPolicy except for a break-glass role, and alerts on any login.

Remember this

  1. 1

    RPO and RTO: Recovery Point Objective is how much data you can lose (time since last good backup). Recovery Time Objective is how long recovery may take. Both come from the business, not from engineers.

  2. 2

    AWS Backup: backup plans (schedule, lifecycle to cold storage, retention) applied to resources by tag (backup=daily). One place for EBS, RDS, Aurora, DynamoDB, EFS, S3, and more. Backup Audit Manager reports on coverage.

  3. 3

    Copies elsewhere: copy backups to a vault in a separate backup account (and often another region). If the prod account is compromised, attackers can't delete those copies.

  4. 4

    Immutability: Backup Vault Lock in compliance mode prevents anyone, including root, from deleting backups before retention ends. S3 Object Lock does the same for objects (governance or compliance mode).

  5. 5

    Restore drills: a backup you haven't restored is a hope. Schedule restores (AWS Backup restore testing can automate this), check data, and time them against RTO.

  6. 6

    Native features still matter: S3 versioning, RDS point-in-time restore, DynamoDB PITR give fine-grained recovery. AWS Backup adds central policy, copies, and immutability.

Explain it without notes

01

Why are multi-AZ replication and backups both needed?

02

What's the purpose of a cross-account, locked backup vault?

Practice

01

Create an AWS Backup plan that backs up resources tagged backup=daily and keeps them 14 days.

02

Run a restore of an EBS volume backup and verify its contents.

Trade-offs

  • ↔

    More frequent backups and longer retention lower RPO and widen recovery options but cost more storage. Vault Lock and Object Lock give strong protection but make mistakes in retention settings permanent. Restore drills cost time and money but are the only proof backups work.

Done when you can

  • RPO and RTO are agreed for each data store.

  • Backups are taken by policy, copied to another account, and locked.

  • I run restore drills and record restore times.