Investigation 028 - Persistent State Principles

The data returned.
Did the participant?

A failed cluster member is replaced. The replacement reconnects to the exact same durable state. Its peers still refuse to accept it.

Begin the investigation down
Member A
recognized
Member B
recognized
Replacement
unrecognized

Author's Note

Persistence, placement, and interface are settled. Continuity is not.

INV-025 gave storage a lifecycle the platform owns independently of any process. INV-026 co-decided the placement of computation and storage under shared topology. INV-027 let any storage backend participate through a stable contract. We inherit all three rather than rediscovering them.

Every one of those contracts assumed that once a replacement process reconnects to surviving data, recovery is complete. A process is an execution — it begins, performs work, and stops. A long-lived logical participant is something else: a distributed system's memory of who leads, who owns a partition, who is allowed to vote. Replacing the execution does not automatically replace the participant.

When a member stops responding, the platform cannot tell whether it has crashed or is merely slow — the same delayed and partial observation INV-011 identified for compute. We assume here that a replacement has already been correctly authorized; preventing two executions from believing they hold the same identity at once is a real problem this investigation acknowledges and defers.

Possessing the same data does not establish that the same participant has returned.

Prologue

The replacement has everything. The cluster still refuses it.

A distributed database runs three replicas. One fails unexpectedly. The platform notices, creates a replacement, and reconnects it to the surviving volume. Every byte is intact. From the platform's perspective, recovery is complete.

Foundation

Persistent storage survives process failure and reattaches automatically.

Assumption

Data surviving is sufficient for the cluster to accept the replacement.

Incident

The remaining replicas still expect the member that disappeared. They do not recognize the newcomer.

Mystery

If nothing about storage, placement, or the interface has failed, what has the platform not yet provided?

First Principles

A process executes. A participant endures.

A process begins, performs work, and stops; the platform can destroy it and create another. A distributed database's replica does more than execute — it participates. The cluster remembers who joined, who leads, and who owns particular responsibilities, and those memories attach to a participant, not to whichever process happens to be running.

Interchangeable

Web servers, batch workers

One instance disappears, another appears, and traffic continues without anyone caring which execution answered.

Not interchangeable

Replicated databases, consensus clusters

A notebook a newcomer carries may hold every decision recorded so far. It does not make the newcomer the participant who left the room.

Deferred question

Slow, or gone?

The platform cannot distinguish a crashed member from a slow one. We assume a failed member has been safely identified; preventing two executions from sharing one identity is left to a later investigation.

Restoring bytes does not restore the relationships a cluster remembers.

Naive Architecture

Any replacement. Same reconnected state.

The platform owns execution; the application owns meaning. When a process fails, start another and let it reconnect to the same durable volume. Whether it receives the same name, the same address, or is recognized by its peers is not the platform's concern.

Failed Processexecution stops
Anonymous Replacementnew name, new address
Reconnected Volumesame durable bytes

The Architecture That Almost Worked

Same data. Equivalent execution. No application ever notices.

The replacement starts with exactly the files its predecessor left behind. No data is copied, nothing is recreated. For batch processors, web servers, and background workers — anything that only needs equivalent work to continue — this is not merely acceptable, it is ideal.

What it preserves

Every previous contract

Data survives, placement respects storage topology, and any storage backend integrates through the same interface.

What it assumes

Interchangeable workers

If the replacement has the same data, nothing else about the system should care who is running it.

The architecture fractures only when the replacement tries to rejoin peers who remember someone specific.

Breaking Our Design

Three independent pressures test anonymous replacement.

Each experiment starts from its own clean state and stops at its assigned discovery.

EPISODE 01

The Name That Changes

Pressure. A member fails mid-operation. Its durable state is intact and owned. The platform replaces it anonymously, under a fresh name.

Prediction. If the data is correct, rejoining the cluster should be a formality.

Experiment

Establish the original member and its owned durable state, fail its execution, replace it anonymously, then attempt to rejoin.

Observation

The replacement carries the correct data under a different name.

Failure

The remaining peers still consider the original participant missing and reject the newcomer.

Discovery

Persistence is not participation. If peers require the same participant name to accept a return, the platform must guarantee that stable name.

Next pressure

Even with a stable name, can every peer still find where it currently runs?

Boundary

Discoverability of the current address is not tested here.

Fail an owned member, replace it anonymously, attempt rejoin.

Member namenone
Owned durable statenone
Executionnot started
Rejoin attemptnot attempted

Prediction: if the data is correct, rejoining should be a formality. Start the member to find out.

EPISODE 02

The Network Identity That Disappears

Pressure. The platform now guarantees a stable name for every replacement. This replacement starts on a different machine than its predecessor.

Prediction. A stable name should be enough for peers to find the replacement wherever it runs.

Experiment

Move the stably-named replacement to a new address, send to its last-known address, then refresh the platform-derived mapping.

Observation

Messages sent to the old address never arrive; the name alone does not carry the current location.

Failure

A peer cannot tell whether the target is unreachable or mid-recovery while the mapping is stale.

Discovery

A stable name must always resolve to the current address of its bearer; observations and mappings of that address may themselves be delayed.

Next pressure

With identity and discoverability both guaranteed, does replacement order still matter?

Boundary

Recovery order and sequencing are not tested here.

Move a stably-named member, then resolve its current address.

Stable nameguaranteed
Current addressoriginal machine
Message to last-known addressnot sent
Platform-derived mappingstale

Prediction: a stable name should be enough. Move the member to test that.

EPISODE 03

The Order That Cannot Be Assumed

Pressure. A three-member quorum cluster needs to bootstrap. Identity and discoverability are already guaranteed for every member.

Prediction. With identity and discoverability solved, replacing all three members at once should recover the cluster fastest.

Experiment

Attempt concurrent recovery of all three quorum members at once, then observe the bootstrap protocol.

Observation

No member has an established peer to bootstrap against, and quorum drops to zero before any replacement establishes itself.

Failure

The protocol's prerequisite for a sequenced start is violated; the cluster cannot confirm any replacement or make progress.

Discovery

When protocol correctness depends on sequence, recovery order is an architectural constraint, not a scheduling preference.

Next question

What replaces anonymous, unordered recovery as the platform's contract? Chapter 07 formalizes the answer once all three experiments are complete.

Boundary

The specific sequencing mechanism is not derived here.

Replace a 3-member quorum concurrently and observe the bootstrap.

Member Arunning
Member Brunning
Member Crunning
Quorum3 of 3

Prediction: replacing all three at once should recover the cluster fastest. Test it.

Review the three experiments
  1. Persistence is not participation; a replacement that returns under a new name is not recognized as the participant that disappeared.
  2. A stable name only restores communication once it always resolves to the current address of its bearer.
  3. When protocol correctness depends on sequence, recovery order becomes an architectural constraint rather than a scheduling choice.

The Turning Point

Stop asking what survives.
Ask who returns.

Logical Participantlong-lived responsibility
Identity Continuityplatform owned
Replacement Executioncurrent process
A process may fail. The logical participant must endure.

The Stable Identity Contract

One boundary. Four responsibilities.

Contract not yet earned. Complete all three experiments before formalizing the contract.

Only Now: Kubernetes

StatefulSet, its governing Service, and per-Pod claims realize this contract.

Kubernetes realizes the identity continuity contract through StatefulSet, working together with a governing headless Service and per-Pod storage claims. Each replica receives a stable ordinal Pod name and a stable network identity, and the governing Service provides the DNS domain through which peers perform current discovery at a high level. Depending on how the workload is configured, Pods can start and terminate in an ordered, one-at-a-time sequence when the application requires it. volumeClaimTemplates give each ordinal Pod its own persistent volume claim rather than a claim shared across replicas.

Identity

Stable ordinal names

Each Pod keeps its ordinal name and network identity across replacement.

Discovery

Governing Service

A headless Service gives the set its DNS domain for locating each member's current address.

Storage

Per-Pod claims

volumeClaimTemplates allocate a distinct persistent volume claim to each ordinal Pod.

StatefulSet does not make failure detection certain, does not fence duplicate identities from both believing they hold the same role, does not guarantee an application will accept a returning member, does not make DNS resolution instant or always current, and does not guarantee exclusive backend access to a volume. Ordered, one-at-a-time Pod management is a configurable default, not a universal guarantee — Pod management policy may permit parallel Pod management when the application allows it.

Engineering Reflection

Timeless Engineering Principle

Persistence answers what survives failure. Identity answers who returns after failure. A platform that preserves one without the other can recover data while still failing to recover the distributed system that depends on it.

Architectural Honesty

Keep anonymous replacement

Stateless services, batch processing, and interchangeable workers care that work continues, not which execution performs it. Identity continuity adds cost without improving correctness here.

Introduce identity continuity

Replicated databases, consensus systems, and message brokers assign durable responsibilities to individual members. Continuity becomes necessary, not optional, for these.

Costs Accepted

Reduced interchangeability

Replacements can no longer be treated as anonymous workers.

Sequenced recovery

Correct ordering may take precedence over maximum recovery parallelism.

Identity metadata

The platform must track participant continuity, not merely count running processes.

New failure modes

Replacing a participant too early risks two executions that both believe they hold the same identity.

Investigation Exercise

Predict, then trace, a replacement against a membership-aware cluster.

Prediction

A database member returns with the same data but a different name. Predict whether its peers should accept it as the missing participant.

Experiment

Compare that replacement with one that also preserves its name, its current discoverable location, and any required recovery order.

Observation

Notice that the bytes stored on disk never decided whether the cluster trusted the returning execution.

Reflection

Ask what a platform must additionally guarantee once an application assigns durable responsibilities to named participants.

o o o
Run the synthesis trace after writing your prediction.

Bridge to INV-029

Identity establishes who returns.
Not what it is meant to do.

The platform no longer merely reconnects a replacement to surviving data.

It preserves a stable participant identity across process replacement.

It preserves discoverability of that identity's current execution.

It preserves the recovery sequence some distributed protocols require.

Kubernetes realizes this contract through StatefulSet, its governing Service, and per-Pod storage claims — yet the returning participant still may not know its intended role.

A logical member returns with stable identity, current discoverability, ordered recovery, and its own durable state, while its intended role remains unresolved

Next Investigation

INV-029 - The Configuration Problem

Identity tells the cluster who has returned. It does not tell that participant how to behave.

Intellectual Lineage

This investigation inherits desired-state reconciliation, explicit ownership, and architectural decomposition from earlier investigations. Raft and Paxos provide source-grounded examples of distributed protocols whose safety and progress depend on remembered participants and quorum membership rather than anonymous process replacement. Kubernetes StatefulSet is one realization of the broader continuity contract derived here; this investigation does not claim that StatefulSet invented stable participant identity.

Deliberate Simplifications Ledger

Safe stateful-member replacement and prevention of two executions holding one identity simultaneouslyFuture: Stateful Member Fencing (unassigned)
Mechanisms for stable participant naming and always-discoverable current addressesFuture: Stateful Identity Realization (unassigned)
Mechanisms for ordered lifecycle and per-participant storage allocationFuture: Stateful Lifecycle Realization (unassigned)
Exclusive ownership enforcement for per-participant volumesFuture: Stateful Storage Access Control (unassigned)