One agreed history
INV-008 proved consensus is unavoidable. Every controller now builds on the same accepted past.
Controller A reads a Deployment with three replicas and decides to scale to five. Controller B reads the same Deployment and decides to update its labels. Both are correct. Both are accepted. Inspect the Deployment: labels updated, replicas still three. Controller A's change has vanished.
Nobody failed. Nobody made a mistake. The history remained singular throughout. And yet a writer's work was silently erased.
Independent machines, separated by an unreliable network, now agree on one history. Every controller reads from the same accepted past. The control plane finally has the foundation it needs.
Now consider two controllers, both operating correctly. Both read the same Deployment at approximately the same moment — three replicas. Controller A decides to scale to five. Controller B decides to update its labels. Both observations are correct. Both decisions are valid. Both changes are accepted.
Inspect the Deployment. The labels have been updated. The replica count is still three. Controller A's scaling change has vanished. Nobody failed. Nobody made a mistake. The history remained singular throughout. And yet a writer's work was silently erased.
This is not a consensus problem. The question is not whether machines agree on history. The question is whether a writer is still modifying the same version of history that everyone else is building upon.
Before searching for an architecture, we must accept the conditions it must survive.
Controllers, schedulers, operators, and tools all read and modify cluster state concurrently. No coordination prevents two participants from deciding to modify the same object.
When a write is accepted, the shared history advances. Participants who haven't re-read it are now working from an older version of truth — with no notification.
There is always a gap between reading and writing. In a busy system, that gap is constant, not an edge case.
The question is not whether concurrent modification happens. The question is whether the system can detect — and safely handle — the moment a writer's understanding of reality has grown stale.
Consensus already ensures the shared history never diverges. Every accepted write becomes part of one authoritative sequence. Surely concurrent writers simply have their changes recorded in order, and both contributions survive.
No conflicts are detected. No errors are reported. Both operations succeed. Surely nothing is wrong.
Overwritten, without warning, by a writer who had read an older version and submitted a complete replacement of it.
Ready: Both controllers read the same state, then each submits an individually valid change.
We will not allow two writers to modify the same object at the same time. When Controller A begins modifying a Deployment, the system places a lock on it. Controller B's attempt is rejected until the lock is released.
For a moment, the architecture feels correct. It is — until a harder question arrives. What happens when Controller A acquires the lock and then fails? Its process crashes, its connection drops, it restarts and loses all memory of the lock it was holding. The lock is still held. Nobody can modify the Deployment.
Set the timeout too short, and a slow controller loses work it was genuinely completing. Set it too long, and a failed controller blocks the entire object for minutes. There is no safe value — every timeout is a guess.
A locking mechanism that stops working when participants fail is not a solution. It is the original problem in a different form.
The architecture we proposed — accept all writes, consensus keeps history singular — was exactly what any engineer would build first. It is not unreasonable. It has a hole in it.
Two controllers read the same history. Both changes are valid. Both are accepted. Yet one silently replaces the other — not rejected, not rolled back, simply overwritten by a writer who read an older version.
Agreement protects the past. It does not automatically protect the future.A controller's local view tells it what reality was, not what reality is now. Controller B still believes it's acting on the latest reality, unaware Controller A already advanced the shared history.
Consensus answers "what is the accepted history?" This is a different question: "is the history I observed still the accepted history?"If every accepted history looked identical — no identifiers, no generations — the system could never know whether a writer's understanding was outdated. The problem is not concurrency. The problem is identity.
Every accepted history must be distinguishable from every history that came before it.Kubernetes expresses the versioning contract through a field: ResourceVersion. When a controller submits an update, Kubernetes compares the version it observed with the version currently stored. Mismatch means the history has moved — the update is rejected, not silently accepted.
The name is unimportant. Any distributed control plane solving this problem requires an equivalent concept.Ready: Both controllers observed the same version. Watch what happens when the second writer's version has gone stale.
Consensus constructs one shared history. Versioning protects that shared history as it evolves.
The contradiction becomes surprisingly simple once framed correctly. If the system can compare the history a writer originally observed with the history currently accepted, it can tell whether a writer is still editing the world it examined. If they match, the write proceeds. If they differ, the writer must first observe the latest reality before deciding again.
This is not an optimization. It is the only way to ensure that independently correct participants cannot unknowingly destroy one another's work.
ResourceVersion is not merely metadata. It is Kubernetes' way of identifying which accepted version of reality an object represents — the concrete expression of an architectural contract every distributed control plane must satisfy.
Consensus solved the first. Versioned truth solves the second. Together they establish this engineering contract.
Independent replicas are insufficient. Agreement must precede acceptance.
A shared history that evolves over time must give every accepted state a unique identity.
The responsibility for detecting stale understanding belongs to the system, never the writer's assumptions.
Never silently replace newer accepted history with updates derived from older observations. Correctness precedes convenience.
A rejected update returns to reconciliation — the writer observes latest reality before deciding again.
Generation numbers, revision identifiers, version vectors — the implementation is replaceable. The responsibility is not.
ResourceVersion satisfies two distinct debts. Used for optimistic concurrency (this investigation), a writer proves it is modifying the current version. Used for watch resumption (the debt from INV-004/005), a watcher uses it as a position marker to resume a broken stream. Same field, two contracts — both made possible by one property: every accepted change carries a unique, monotonically advancing identity.
"What is the accepted history?" and "Am I still modifying that accepted history?" are fundamentally different questions. A Git push rejected because the remote moved forward, a collaborative document merging concurrent edits, a database transaction using compare-and-swap — the same architectural pattern appears wherever independent actors cooperate on shared state.
A single-writer system, or one where all modification flows through one serializing component, never produces a lost update. Versioned truth is necessary only when concurrent writers, a read-then-write delay, and invisible-until-inspected corruption all coexist — exactly Kubernetes' conditions.
A writer must observe current state before modifying it; a race means rejection and re-fetch. High-contention objects may need multiple retries. That latency is the deliberate cost of: no writer silently destroys another writer's work.
Step 1: kubectl create deployment nginx --image=nginx, then kubectl get deployment nginx -o yaml — note metadata.resourceVersion.
Step 2: kubectl scale deployment nginx --replicas=5, then re-inspect. The ResourceVersion has changed, even though it's the same logical object.
Step 3: Export the Deployment to a file. Modify it again from another terminal. Compare the file's resourceVersion against the live one — your file now represents an older version of reality.
Reflect: ResourceVersion identifies which accepted version of reality an object represents, letting independent participants detect stale understanding before silently overwriting each other's work.
Hypothesis: Once every machine agrees on one shared history, concurrent updates should always be safe.
Experiment: Controller A changes replica count and updates the shared history first. Controller B then submits its label update using the older view it observed.
Observe: Both were individually valid. One silently replaced the other. Agreement about the current history did not guarantee safe creation of the next history.
Hypothesis: A system can safely detect stale updates only if every accepted history can be distinguished from the previous one.
Experiment: Repeat the scenario, but every accepted history now receives a version identifier compared against the writer's observed version before acceptance.
Observe: No update is silently lost. Writers either modify the latest history, or reconcile first. Correctness and progress are both protected.
The control plane now has two complementary guarantees: consensus protects shared history from diverging across machines, and versioned truth protects it from being silently overwritten by concurrent writers. The architecture appears complete. It is not.
Multiple instances of the same controller manager run simultaneously for resilience. Every instance observes the same cluster and is capable of making the correct decision. Versioned truth prevents corruption if two instances both attempt the same reconciliation — but it does not answer whether both should have attempted the work in the first place.