Ephemeral state — study guide

The concept's fragments, read in order.

Gone when it restarts

Some of what a running system holds is written in disappearing ink. A process keeps its working data in memory, and that memory is gone the instant the process exits. A container carries its own scratch space, and that space is wiped when the container is removed and started again from its image. A virtual machine can be rolled back to an earlier snapshot and lose everything written since. This vanishing state is called ephemeral state, and the trouble starts when you assume it will still be there and it is not.

The split that runs through everything here is between state that may safely disappear and state that must be made to last. Scratch files, caches, and the in-memory bookkeeping a program builds up while it runs are fine to lose, because they can be rebuilt. A user's uploaded photo, the rows in a database, the one log you actually needed — those must survive a restart, and survival is something you arrange on purpose, not something you get for free.

Everything here is how to tell the two apart and how to keep the second kind: why a container's writable space is temporary, why that is often a feature rather than a bug, how a volume or an external store holds data outside the thing that keeps restarting, and why the same trap waits in virtual machines. Get the split wrong and your work is gone; get it right and you can throw the compute away and lose nothing.

The writable layer is temporary

container (first run) writable layer your file read-only image layer read-only image layer read-only image layer podman rm, then podman run recreates from the image container (recreated) writable layer empty: file gone read-only image layer read-only image layer read-only image layer the read-only image survives; the writable layer, and your file, do not
A file written to a container's writable layer, on top of read-only image layers, is gone when the container is removed and recreated from the same image.

A container runs from an image, and that image is a stack of read-only layers that never change. When the container needs to write something, it cannot touch those layers, so it writes to a thin writable layer stacked on top, unique to that one container. Every file the running program creates or modifies lands there. The catch is in the word temporary: that writable layer is discarded when the container is removed, so a fresh container recreated from the same image starts clean, with nothing you wrote last time. like a hotel room wiped clean for the next guest: anything you leave inside is gone once you check out, so you do not store there what you need to keep

This is not a bug or a crash you can avoid by being careful. It is how the layering works by design. The read-only image is shared and reproducible; the writable layer is per-container scratch space that dies with the container, and it cannot easily be pulled out to another container either. Docker works the same way; podman is the tool we use here. Treat anything living only in that layer as already lost, because a podman rm followed by a fresh podman run will prove it.

A process behaves the same way with its memory: the data a program keeps in RAM while it runs is gone the moment the process ends. The pattern is one and the same — the live thing owns some fast, private, disappearing storage, and when the live thing goes away, so does the storage. Knowing exactly which storage is disappearing is the whole game.

Disposable on purpose

Losing the writable layer sounds like a defect until you see what it buys. When every container is rebuilt from a known image and changed by nothing after that, there is nothing for it to drift into. Configuration drift — the slow divergence where one long-lived machine accumulates hand-edits and patches until it is subtly unlike its neighbors — simply cannot happen to a thing that is never edited in place. You do not fix a misbehaving instance; you delete it and let a clean one take its place. This is the idea behind immutable infrastructure.

Because every instance is rebuilt from the same image, instances become interchangeable, and interchangeable instances are cheap to replace and to multiply. Need more capacity? Start more identical copies. One is misbehaving? Kill it and start another; the replacement is the same as what it replaced. The industry shorthand is cattle, not pets: you do not name a disposable instance and nurse it back to health, you cull it and move on.

None of that works if the instance is secretly hoarding something you need. The whole benefit rests on a promise the operator has to keep — that everything worth keeping lives somewhere other than the disposable compute. Ephemerality is a feature only when nothing important is trusted to the ephemeral part.

Mounting something that lasts

container (first run) podman rm, then podman run again container (recreated) mount with -v re-attach, data intact named volume data persists outside the container
A named volume lives outside the container lifecycle, so when the container is removed and recreated the volume and its data persist and re-attach to the new container.

If the writable layer dies with the container, the fix is to keep the data somewhere that is not the container. That somewhere gets mounted in, so the program reads and writes a normal-looking directory while the bytes actually live outside the container's lifecycle. like a safe-deposit box kept outside the room: it is not part of the room, so it survives when the room is cleared and reset There are two common shapes for this, and podman offers both through the -v flag.

A named volume is storage podman manages for you, given a name so it outlives any one container. You attach it with podman run -v mydata:/var/lib/app, and podman creates the volume if it does not already exist. Crucially, a named volume is not deleted when the container is removed — not by podman rm, and not even by the volume-clearing options — so the next container you start can mount the same podman volume and find the data waiting. A bind mount is the other shape: podman run -v /host/path:/var/lib/app maps a directory from the host straight into the container, so the data plainly lives on the host and the container is just borrowing a view of it.

Either way, the container goes back to being disposable. Remove it, recreate it, and the writable layer is thrown away as always — but the volume, being outside that lifecycle, re-attaches to the new container intact. The rule is simple: what you mount survives; what you leave in the writable layer does not.

Keep state out of the compute

stateless container stateless container stateless container read and write state external state store database or object bucket any container is replaceable
Several interchangeable stateless containers all read and write one external state store, so any container can be replaced without losing state.

A volume keeps data alive across restarts, but it still ties that data to a particular place. The sturdier pattern moves the real state further out entirely: into a backing service built to hold it — a database for structured records, an object bucket for files and blobs. The container keeps nothing durable of its own. It reads what it needs from the store and writes what must last back to the store, and that is all the state it touches.

This is the Twelve-Factor rule of thumb: processes should be stateless and share nothing, and any data that has to persist belongs in a stateful backing service rather than in the process. The payoff is that the compute becomes genuinely interchangeable. like interchangeable rental cars: any one will do because nothing of yours is stored in it, and your belongings travel separately with you Run one container or run ten identical ones, all pointed at the same database or bucket; any of them can serve a request, and any of them can be killed mid-life without taking state down with it, because none of them was holding the state in the first place.

That is the difference between merely persisting data and designing so the compute never mattered. A volume answers "how do I not lose this when the container restarts"; an external store answers "how do I stop caring which container, or how many, are even running". The state lives in one durable place, and the containers become a crowd of replaceable readers and writers around it.

Virtual machines lose it too

It would be comforting to think this is a container quirk, but the same trap waits one level down, in virtual machines. A VM can be captured as a snapshot — a frozen picture of its disk and state at one moment — and rolled back to it later. Everything written after that snapshot was taken is gone in the rollback, exactly the way a fresh container loses the old writable layer. The mechanism differs; the lesson does not.

Ephemeral cloud instances make the point sharper. An instance can be given temporary block storage that lives only as long as that instance in a particular running state: the data on it survives a plain reboot, but it does not survive the instance being stopped, hibernated, or terminated. The provider is blunt about it — do not keep anything valuable there, copy what matters to durable storage. The storage is even named for what it is, with device names built from the word ephemeral.

So the split is not about containers versus VMs. It is about which storage is attached to the disposable thing and which storage is kept apart from it. Whether the disposable thing is a process, a container, or a whole virtual machine, the discipline is identical: know what disappears when it resets, and put anything you need to keep somewhere that does not reset with it.

Designing so restarts are safe

Stateless design is the habit that makes all of this safe to rely on. A stateless instance holds no durable state of its own: it takes its settings from its environment when it starts, does its work, and sends anything that must last to an external store. Nothing important accumulates inside it, so nothing important is riding on it when it stops. Killing and replacing an instance becomes a routine event rather than a data-loss incident.

The test is a plain one. Pick any running instance and ask what breaks if it dies right now, with no warning. If the honest answer is "some data no one else has," it is not stateless, and that data is one crash away from gone. If the answer is "nothing that a fresh instance cannot pick up from the store," you have arrived — the compute is disposable and the state is safe, and those two facts are the same fact seen from two sides.

This is why stateless services scale and heal so cleanly. Load balancers can spread requests across any instance because they are all equivalent. A crashed instance is simply replaced, and the replacement reads the current state from the store and carries on. You get replaceability and resilience for the price of one rule kept honestly: keep the durable state out of the disposable compute.

Why your agent forgot everything

Here is where this stops being theory. You build an agent that does real work — reads a task, gathers notes, writes files as it goes — and it runs inside a container. It finishes a run, you start it again the next day, and everything it learned is gone. It did not malfunction. It wrote its notes to the writable layer, the container was recreated between runs, and the layer went with the old container. The agent forgot because you stored its memory in disappearing ink.

The fix is to persist what matters and treat the compute as disposable. Anything the agent must remember across runs goes to a mounted volume or, better, to an external store the next run can read: a database of what it has done, a bucket of the artifacts it produced. A homelab service has the same shape — rebuild the container to update it and the config or data left inside is wiped, unless it was mounted or kept in a store that outlives the container. In both cases the compute is meant to be disposable; only the state you deliberately put outside it survives.

This becomes load-bearing the moment more than one disposable thing is in play — agent pipelines and orchestration systems that start, stop, and replace containers constantly. The single question that keeps them honest is the one from every fragment here: when this thing restarts or gets replaced, what disappears, and did I put everything I need to keep somewhere that does not disappear with it?