Backup strategy — study guide

The concept's fragments, read in order.

A copy you can actually go back to

A backup is a separate copy of your data, kept so you can restore it after the original is lost. That is the whole idea, and the word that matters is separate: the copy has to survive the thing that destroyed the first one, or it was never a backup at all.

The failures a backup answers are the ones that redundancy cannot. RAID and similar arrangements keep you running through a dead disk, but they keep every copy in lockstep with the live data, so a deletion, a corruption, or a ransomware sweep hits all of them together, and a fire or a theft takes the whole machine at once. Redundancy is not a backup, precisely because its copies share the fate of the original.

Everything here is how to keep a copy that does not share that fate: what actually counts as a backup, how full and incremental copies trade time against convenience, the 3-2-1 rule for spreading copies around, the choice between a disk on the shelf and a bucket in the cloud, keeping past versions, and the one step people skip until the day they cannot afford to.

Separate, or it is not a backup

A real backup is independent of the original. It lives on a different medium, ideally in a different place, and nothing that happens to the live data reaches it automatically. Deleting or corrupting the original leaves the backup untouched, because the backup is a point-in-time copy that exists on its own, not a reflection of the current state.

This is exactly where a live-synced copy fails the test. Sync keeps two locations identical in real time, which is convenient and feels safe, but it means a deletion on one side is faithfully deleted on the other, and a corruption copies straight through. A mirror that propagates your mistakes is not a backup; it is a second victim. The same is true of a second copy on the same array, which shares the machine's fate down to the last disk.

The line to hold is simple. If the copy changes the moment the original changes, it cannot save you from a change you did not want. A backup earns the name only by being far enough out of band that the disaster stops before it gets there.

Copy everything, or just what changed

one full backup, then a chain of incrementals full all data inc 1 changes inc 2 changes inc 3 changes inc 4 changes each incremental = only what changed since the previous backup chosen point restore = full + every incremental up to the chosen point a missing or corrupt link in the chain breaks the restore time
A restore rebuilds the data from the full backup plus every incremental in the chain up to the chosen point in time.

A full backup copies all of the data, every time it runs. It is the simplest kind to reason about and to restore from, because one full backup is a complete standalone copy, but it is also the slowest and largest to make, since it repeats work on everything whether or not anything changed.

An incremental backup copies only what has changed since the last backup of any kind. Each run is small and fast, because it records just the difference, but that thrift has a cost at restore time: recovering means starting from the full backup and then applying every incremental in the chain, in order, up to the point you want. If any link in that chain is missing or corrupt, the recovery can fail. A differential backup sits between the two — it copies everything changed since the last full backup, so it grows over time between fulls, and a restore needs only the full plus the single latest differential.

The choice is a trade between the cost of making copies and the cost of restoring them. Incrementals are cheap to write and dear to restore; differentials cost more to write and less to restore; a full is dear both ways but depends on nothing else. Most schemes mix them: a periodic full to anchor the chain, and lighter copies in between.

Three, two, one

3 copies, 2 media types, 1 offsite on-site copy 1 media: disk copy 2 media: tape / NAS 2 different media types a different location copy 3 offsite bucket 1 copy offsite three copies in total, so no single failure takes them all
The 3-2-1 rule: keep three copies of the data, spread across two different media types, with one copy kept offsite.

The 3-2-1 rule is the standard shorthand for a backup that can actually survive trouble. Keep three copies of the data, on two different kinds of media, with at least one copy offsite. Each number closes a gap the others leave open.

Three copies means the original plus two backups, so that losing any one — or even two at once — still leaves something to restore from. Two different media types guards against a whole class of failure at once: if both backups sit on the same kind of drive from the same batch, a single defect or a single vulnerability can take them together. One offsite copy is the part that survives a local disaster; a fire, a flood, or a theft that takes the building takes every copy inside it, and only the copy somewhere else lives through it.

The rule is deliberately blunt, and that is its strength — it is easy to remember and easy to check. Modern variants extend it for ransomware, adding an immutable or air-gapped copy, but those are additions to 3-2-1, not replacements for it.

A drive on the shelf, or a bucket in the cloud

your site machine the original local disk fast, on-hand one fire, flood, or theft takes the machine and disk together send copies over the network restore: needs network, costs to retrieve offsite bucket remote object storage survives local disaster local disk shares the site's fate; the offsite bucket does not
A local disk sits beside the machine and shares the site's fate in a fire or theft, while an offsite bucket reached over the network survives the local disaster.

Once you accept that at least one copy belongs offsite, the practical choice is where the other copies live. A local disk backup is a drive kept near the machine — an external drive, or a second drive in the same room. It is fast to write to and fast to restore from, and it is right there when you need it. Its weakness is exactly its convenience: it sits inside the same building as the original, so the fire, flood, or theft that takes one takes both. It shares the site's fate.

The offsite alternative is a bucket. A bucket is remote object storage — a service, usually reached over the internet, that holds copies you send to it and hands them back when you ask, addressed by name rather than by disk location. Because it lives somewhere else entirely, a bucket survives the local disaster that wipes out everything on-premises like a spare key kept at a neighbor's house rather than under your own doormat: it is still reachable when the disaster that hits your home takes everything inside it. The trade is that reaching it depends on the network, and pulling a large restore back down takes time and often costs money, since providers charge to retrieve and move data out.

So the two are not rivals so much as roles. The local disk is your fast, everyday recovery for the ordinary mistakes; the offsite bucket is the copy that is still there when the everyday copy is gone with the building. A sound plan keeps both, which is one of the reasons the 3-2-1 rule asks for a copy offsite in the first place.

How far back can you reach

Retention and versioning decide how many past copies you keep, and for how long. A backup that only ever holds the most recent state answers one question well — get me back to how things were a moment ago — but it is helpless against the damage you do not notice right away. If a file has been quietly corrupted, or a folder encrypted by ransomware, the latest backup may faithfully hold the broken version, and the good one is already gone.

Versioning fixes this by keeping point-in-time copies rather than a single overwriting one, so you can step back to a moment before the damage rather than only to the latest, already-broken state like an undo history: you can step back to a moment before the damage was done, not just to the latest, already-broken state. Retention is the companion policy: it sets how long those past versions are kept before they age out, balancing the reach you want against the storage it costs. Keep too few and the corruption outlives your oldest good copy; keep everything forever and you pay for it.

The payoff is a plan that survives slow disasters as well as sudden ones. Some failures announce themselves; others sit in the data for weeks. Only a history of versions lets you recover from the ones you find late.

An untested backup is a hope

A backup only counts once you have restored from it. Until then, its existence proves that a copy was written, not that the copy is complete, readable, and able to bring the data back — and those are the things you actually need on the worst day. An untested backup is like a fire escape you have never walked: you only discover it is blocked when it is far too late like a fire escape you have never walked: you only find out whether it is blocked at the moment you most need it to work. So test restores, on a schedule, before you are relying on them for real.

Testing also forces two decisions you should make while calm rather than in a crisis. One is your recovery point objective, or RPO: the point in time you must be able to recover to, which is really a statement of how much recent data you can afford to lose. The other is your recovery time objective, or RTO: how quickly you must be back up and running after a loss. The names sound alike, but one is about lost data and the other is about lost time.

Those two numbers turn a vague wish into a spec you can test against. An incremental chain that restores correctly but takes a day to replay may still fail an RTO of an hour, and a nightly backup cannot meet an RPO of five minutes. You find that out by restoring, not by hoping.

The plan that runs itself

Backup strategy earns its keep the moment you run anything you would hate to lose. A game server accumulates worlds, saves, and player state; a homelab holds configuration, media, and the small pile of files you would never be able to reconstruct. None of that is protected by the fact that it lives on a healthy disk, because a healthy disk still deletes what you tell it to and still corrupts what a bad write corrupts.

The plan that actually protects it has a shape by now. It runs automatically, because a backup that depends on you remembering is a backup that eventually stops happening. It follows 3-2-1 for the data you cannot recreate — three copies, two media types, one offsite in a bucket — so no single disaster takes everything. It keeps enough versions to reach back past a corruption you notice late. And it has been restore-tested, so the copies are known to bring the data back rather than merely hoped to.

The point of all of it is a boring outcome: when a disk dies, a file is deleted, or the building floods, you restore and move on. The work is done up front, quietly and on a schedule, so the bad day is uneventful.