Block versus object storage — study guide
The concept's fragments, read in order.
Two very different shapes of storage
Once you outgrow the disk bolted inside a single machine, storage at scale arrives in two dominant shapes, and they are not variations on one idea. Block storage hands you a raw device: a run of fixed-size blocks with no meaning of their own, which you format with a filesystem and use like a local drive. Object storage hands you a bucket: a flat collection of whole objects, each stored with its metadata under a unique key and fetched over a network API.
The split runs deep because the two answer a different question. Block storage answers "give me byte number so-many on this device," the same primitive a filesystem already knows how to build names on top of. Object storage answers "give me the object stored under this key," and never exposes byte offsets at all. One is a surface to impose structure on; the other is a keyed store you hand whole things to.
Everything here is that contrast and what follows from it: how each is addressed, how you change what is stored, how far each one scales, and which of your workloads belongs on which. Get the shape wrong and you fight the storage the whole way; get it right and it disappears under the job.
Raw blocks, structure is yours to add
Block storage presents a device as a long row of fixed-size blocks, addressed by their position, and nothing more. The blocks carry no inherent structure: no files, no names, no notion of where one thing ends and the next begins like a blank ruled notebook: uniform empty space onto which you impose whatever structure you need. It is the same primitive a physical disk offers, which is exactly why it is useful — it is raw uniform space, and the structure is yours to impose.
Because the device is structureless, you put a filesystem on top of it before it is good for much. The filesystem decides which blocks hold which file's bytes and maps the directory tree of names onto those blocks, turning numbered slots into the named files you actually work with. In the cloud this looks like a volume you attach to one machine and then use in the same way you would use a local hard drive — format it, mount it, store files, install programs.
That attachment is part of the shape. In the classic model a block volume is attached to a single machine at a time, which is what lets that machine treat it as its own private disk with fast, direct reads and writes. Run lsblk on a Linux box and the row of block devices it lists is precisely this layer: raw disks and partitions, sitting underneath whatever filesystem was placed on them.
A whole thing, plus a name
Object storage does not hand you a device to format. It stores each thing you give it as an object: the data itself, a set of metadata describing it, and a unique key that identifies it. You hand over a whole item and get back a way to name it like a coat check, where you hand over a whole item and get a ticket, then present the ticket to get the whole item back; later you present the key and get the whole item back. There is no disk to mount and no offset to seek — the key is the only handle.
Objects live in a bucket, which is simply a container for objects. The bucket is a flat namespace: every object in it is found by its key, not by walking down a tree of folders. The metadata travels with the object as a set of name-value pairs — its type, its size, whatever tags you attach — so the object is a self-describing bundle rather than anonymous bytes.
You reach all of this over the network, not through the operating system's disk layer. Object stores expose an HTTP API: you PUT an object under a key and GET it back by that key, using ordinary HTTP requests against the bucket. That one design choice — a keyed store behind a network API instead of a device on a bus — is what makes object storage behave so differently from a disk.
Offset versus key
The sharpest difference between the two is how you say which data you want. Block storage addresses by offset: the device is a row of numbered blocks, and you name data by its position in that row. Something above — a filesystem, usually — has to remember which offsets hold which file. The address is a number that means nothing except "this far into this device."
Object storage addresses by key. Each object has a unique key within its bucket, and you fetch the object by presenting that key like finding a crate by a unique label in one big open warehouse, rather than by a spot on a numbered shelf you must walk to. You do not navigate to it and you do not know or care where it physically sits; the key is a label that maps straight to the whole object, wherever it lives.
It is tempting to read a key like photos/puppy.jpg as a path through folders, but the namespace is flat. The slashes are just characters in one long key, and the "folders" are a convenience: tools group objects by a common key prefix so a bucket can be browsed as if it had directories. Underneath there is no tree to descend — only a single flat space of keys, each pointing at one object.
Edit in place, or replace whole
How you change what is stored splits the same way the addressing does. Because a block device is a row of addressable blocks attached to one machine, that machine can read and write any block in place, at random, whenever it likes. Changing a few bytes deep inside a large file means rewriting just the blocks those bytes live in; nothing else is touched. This random in-place access is what a database or a running program expects from its disk.
Object storage offers no such thing. There is no key you can seek partway into and overwrite a stretch of bytes — an object is replaced as a whole. To change it, you PUT a new version under the same key, and the store swaps the old object for the new one atomically: a reader gets either the old object or the new one, never a half-written mix. You edit by replacing, not by patching in place.
What object storage gives up in fine-grained editing it makes back in reach. Where a block volume serves its one attached host, a bucket sits behind a network API that enormous numbers of clients can hit at once, each fetching whole objects by key over HTTP. The trade is deliberate: give up in-place edits, gain massively parallel access from anywhere.
Bounded by a device, or spread across many
A block volume is bounded by the device behind it. It has a size, and that size is a hard ceiling: fill the volume and you are out of room until you grow or add hardware. Its data may be replicated across several servers so a failure does not lose it, but the thing you address is still one device with one capacity, used by its one attached machine.
Object storage is built the other way around. A bucket is not a device; it is a flat namespace whose objects are spread across many machines, and it will hold as many objects as you keep adding — there is no single disk whose size limits it. Redundancy is built in rather than bolted on: the store keeps copies of each object across multiple servers, and across separated failure zones, so the loss of any one machine does not lose the data.
That is why the flat namespace matters beyond addressing. With no directory tree to walk and no device boundary to respect, a bucket can scatter its objects wherever there is room and capacity and copies both grow by adding more machines, not a bigger disk. Block storage scales up to the limits of a device; object storage scales out across a fleet.
Picking the right one
The choice follows the access pattern, not taste. If a workload needs random in-place reads and writes to a device it treats as its own, it wants block storage. Databases are the clearest case: they constantly rewrite small records scattered through large files, and they expect a fast local disk to do it on. A machine's boot and operating-system disk is the same story — it is a filesystem on a block volume, mounted and written like any local drive.
If instead a workload stores whole items and later fetches them whole, especially from many places at once, it wants object storage. Backups and archives are written once and read rarely, as complete units. Media files and static website assets are uploaded once and served to countless clients by key. Data lakes pile up vast numbers of objects for later analysis. None of these needs to patch bytes in the middle of a file; all of them want cheap, huge, widely reachable storage.
The failure mode is forcing the wrong shape. Run a busy database against object storage and its constant small in-place updates have nowhere to land, since an object is replaced whole; keep terabytes of write-once archives on a block volume and you pay for a bounded device to do a job a bucket does more cheaply and without limit. Match the shape to the pattern and the storage stops being in the way.
The disk under a VM, the bucket for your assets
This distinction stops being abstract the moment you run your own machines. Spin up a virtual machine in a homelab and its virtual disk is block storage: a device the VM sees as a plain drive, formats with a filesystem, and boots from. To the guest it is indistinguishable from a physical disk — the same numbered blocks with a filesystem laid over them. Run lsblk inside that VM and there it is, the block layer under everything the machine does.
The things that machine produces or serves often belong somewhere else entirely. Game assets, disk images, media files, and archived backups are whole items you store once and fetch by name, sometimes from many clients at once — the exact shape object storage is for. They do not want a disk attached to one host; they want a bucket, where each item lives under its key and is pulled back whole over the network.
Knowing which is which saves real grief. There is no command to list a remote object bucket the way lsblk lists your disks, because a bucket is not a device on your machine — it is reached by key over an API, not mounted. Reach for a block device when something needs a disk it can write in place, and a bucket when something needs to be stored whole and served far and wide.