Redundancy — study guide
The concept's fragments, read in order.
More than one disk, on purpose
RAID (a Redundant Array of Independent Disks, and originally of Inexpensive ones) is the practice of merging two or more disks into a single array so the whole behaves better than any one disk alone. The disks can be spinning drives or solid-state; the array presents them to the rest of the machine as one unit of storage.
Combining disks buys two different things, and a given arrangement may aim at either or both. One is redundancy: keep the data intact and the array serving even when a disk dies. The other is performance: spread the work across several disks so reads and writes finish faster than one disk could manage. Which you get depends entirely on how the array is arranged, and the arrangements have names — the RAID levels.
The reason any of this exists is blunt: disks fail. A single disk is a single thing that will, eventually, stop working, and an array is the standard answer to that fact. Everything here is how the common levels make that trade — speed, safety, and wasted capacity against one another — and the one thing an array conspicuously does not give you.
Every disk is a countdown
A disk is a physical device, and physical devices wear out. Spinning drives have moving parts that age; solid-state drives have cells that tire with use. However careful the handling, every disk carries a failure somewhere in its future, and the only real question is when it arrives.
That makes a lone disk a single point of failure: one component whose death takes everything depending on it down with it. When the machine's data lives on exactly one disk, the disk failing and the data being gone are the same event. The filesystem's careful bookkeeping — the blocks, the records mapping names to data — is only as durable as the one device holding it.
Redundancy is the discipline of surviving that failure on purpose. Instead of one disk whose loss is total, you arrange several so that any one of them can die and the array keeps its data and keeps serving. It does not stop disks from failing; nothing does. It changes a failure from a catastrophe into an inconvenience you have time to fix.
Speed, and no safety net
Striping, known as RAID 0, splits each file into segments and spreads them across every disk in the array. Because the disks can then work in parallel — each reading or writing its share at the same time — the array is faster than any single member, and it wastes no capacity on copies, so all of the raw space is usable. It needs a minimum of two disks.
The catch is that striping stores exactly one copy of everything, scattered. There is no redundancy at all: no second copy, no way to reconstruct a missing piece. If any one disk in the array fails, the segments it held are gone, and since every file is spread across all the disks, losing those segments loses the whole array.
So RAID 0 is the arrangement that spends nothing on safety and puts it all into speed and capacity. It is the right choice only where the data is disposable or reproduced elsewhere, because a single dead disk ends it.
Two copies, always
Mirroring, known as RAID 1, writes the same data to two disks at once, so each disk holds a complete copy of everything. It needs a minimum of two disks, and the two carry identical contents at all times.
The payoff is simple survival. If one disk fails, the other still holds the whole of the data, and the array keeps serving from the survivor without losing anything like a spare tire already mounted and turning: one wheel fails and you keep driving without stopping, at the cost of carrying two. Recovery is straightforward because there is nothing to reconstruct — the good disk already is the answer. The price is capacity: two disks store what one disk's worth of data would hold, so a two-disk mirror gives you only about half of the raw space as usable.
That is the trade mirroring makes plainly. You pay for a full second copy and you get a full second copy, which is exactly what keeps the array alive through one disk's death.
Reconstructing what was lost
Parity is the trick that gives you redundancy without paying for a whole second copy. Alongside the data, the array computes and stores extra check information — parity — distributed across the disks. If a disk dies, its contents can be recomputed from the data still on the survivors plus that parity, so nothing is lost like reconstructing a missing piece from a set of clues and the pieces that remain, without having kept a full second copy.
RAID 5 uses a single set of distributed parity. It needs a minimum of three disks, and because one disk's worth of space goes to parity, an array of n disks gives you the capacity of n minus one. It tolerates exactly one disk failure: lose a second disk before the first is dealt with and the array cannot reconstruct what is missing.
RAID 6 extends the idea with double parity — two independent sets of check information — so it can survive two disks failing at once. That protection costs two disks' worth of space, giving an n-disk array the capacity of n minus two, and it needs a minimum of four disks. The pattern is consistent: more parity buys tolerance for more simultaneous failures, at the cost of more capacity spent on it rather than on data.
The dangerous window
When a disk in a parity array fails, the array keeps running in a degraded state, serving data by recomputing the missing pieces on the fly. To return to full health it has to rebuild: a replacement disk is fitted, and the array reconstructs the dead disk's entire contents onto it from the surviving data and parity. For a large disk that reconstruction can take many hours.
Those hours are the dangerous part, because during them the array has spent its redundancy. A RAID 5 array that has lost one disk has no protection left at all until the rebuild finishes. The rebuild also pins the surviving disks at heavy, sustained reads for its whole duration, and those disks tend to share an age, a batch, and a workload with the one that just died — so a second failure is more likely than pure chance would suggest, not less.
There is a subtler failure too. Reconstruction has to read every remaining disk in full, and if the array hits a sector it cannot read — an unrecoverable read error — on one of those survivors, it cannot compute the missing data for that stripe and the rebuild fails. The odds of striking such an error somewhere in a full-surface read grow with disk size, and for today's multi-terabyte drives the risk is far from negligible. This is why tolerating two failures, as RAID 6 does, is worth so much more than tolerating one: the second layer covers exactly the moment the first has been used up.
RAID is not a backup
RAID protects against exactly one thing: a disk failing. That is a hardware fault, and redundancy handles it well. But most of the ways data actually disappears are not hardware faults, and against those an array offers no protection whatsoever, because its copies are kept in lockstep with the live data.
Delete a file and the array faithfully deletes it from every disk. Let corruption or a mistaken overwrite reach the data and every copy corrupts together; let ransomware encrypt it and every copy is encrypted; let fire or theft take the machine and every disk goes at once like several copies kept perfectly in step: change or destroy the original and every copy changes with it, so only a copy kept separately survives the mistake. The array's copies share one fate, because they exist to mirror the current state — which is precisely the state you are trying to escape when something has gone wrong with the data itself.
So an array keeps you running through a dead disk, and a backup lets you recover when the data is wrong. They answer different failures, and one cannot stand in for the other: a second copy on the same array is not a backup. What a backup actually is — a separate, out-of-band copy, and how to keep one well — is the job of backup strategy, the concept that comes next.
Staying up versus staying safe
The place this earns its keep is any machine that is supposed to stay running. On a server or a homelab, a disk dying at three in the morning should be a task for later that week, not an outage and a scramble. An array turns the dead disk into a degraded-but-serving array and a replacement you swap in on your own schedule — the machine never stopped.
The mistake worth avoiding is reading that uptime as safety. Choosing a RAID level is choosing how much disruption a hardware failure causes: RAID 0 for speed and none of the protection, mirroring or parity for survival, more parity for surviving more disks at once. None of those choices does anything about a file you deleted, data that corrupted, or the whole box being lost.
So treat these as two separate jobs and staff both. Redundancy keeps the machine up through the failure of a disk; keeping the data itself safe is a different job that still needs real backups. An array that has never been backed up is one bad command away from being no protection at all.