One authoritative history, promised
INV-007 told us the control plane needs one agreed history rather than many copies. It deferred the mechanism.
A user asks: "Does this Deployment exist?" Machine A says yes. Machine B says yes. Machine C says no. None of them are lying. Each is faithfully reporting exactly what it has seen.
If the machines responsible for remembering history can themselves disagree, how can the system ever know what really happened?
Three control nodes, each watching the same cluster, each remembering the same history. Or at least, that was the design. A user submits: "Create three replicas." One machine records it immediately. A second is still waiting for the message. The third briefly lost contact with the network.
For a few seconds, every machine believed something different. The first was certain the application existed. The second had never seen it. The third believed the request might have failed entirely. None of them were malfunctioning. Each machine was faithfully reporting exactly what it knew.
The ReplicaSet controller asked "How many Pods should exist?" One answer said three. Another said none. Every component was behaving correctly. Every component was following its contract. And yet, together, the control plane had begun drifting apart — because the system no longer agreed on what had already happened.
If the machines responsible for remembering history can themselves disagree, how can the system ever know what really happened?
A single machine has one memory, one timeline, one history. Nobody to disagree with. The moment history is copied onto several machines, truth stops being local.
Imagine three notebooks. Every important event must be written into all three. If every note reaches every notebook immediately, all three remain identical — life is simple. But if the phone rings before the second notebook is updated, and the power goes out before the third, someone asking "what happened today?" gets a different answer depending on which notebook they open.
None of the notebooks are lying. Each faithfully contains everything it has seen. The disagreement comes from something much simpler: they have not all seen the same history yet.
Failures are not the only source of disagreement. Distance is. The simple act of placing history on multiple machines creates the possibility that those machines temporarily know different things.
The most natural idea requires no distributed systems expertise. Instead of trusting one machine, trust several. Just make copies.
Create three replicas — every machine records it. Delete — every machine records that too. Nothing lost, nothing forgotten, nothing inconsistent.
We assumed writing history into multiple places is a single action. Reality offers no such guarantee — networks delay, drop, and reorder messages.
Ready: Write a change, then let the network quietly drop one delivery.
Create three replicas. Every machine records it. Scale to five. Every machine records that too. Machine B crashes — A and C still remember everything, and reconciliation continues without interruption. History has become resilient.
At this point, it is tempting to believe we are finished. We have eliminated the single point of failure. Every machine stores the same sequence of events. This architecture appears remarkably robust. In fact, many distributed systems begin with exactly this idea.
But hidden inside our demonstration is an assumption so natural we almost overlook it: every example quietly depended upon every machine eventually receiving every write. We never questioned it, because nothing had forced us to.
Our demonstration proved that shared history survives machine failures. It never proved that every machine always shares the same history. Those are not the same thing.
We have only tested storage. We have never tested communication, and we have never tested agreement.
Machine A records "Create a Deployment." Machine B records it. The network is interrupted before Machine C hears. Asked "Does this exist?" — A and B say yes, C says no. All three are honest.
The problem is not corrupted data. It is incomplete observation.The network splits. A and B stay connected; C is isolated. Each side keeps accepting requests. Every new request widens the gap — the partition isn't delaying history, it's creating multiple histories.
Both groups may be internally consistent, healthy, and honestly convinced they are correct.Machine C doesn't respond. Seconds pass. Has it crashed, or is it just slow? From A and B's perspective, every explanation looks exactly the same: silence.
A machine that is slow is indistinguishable from a machine that is dead — no algorithm changes that.Ready: Watch how "slow" and "dead" look identical from the outside, until it's too late either way.
Replication protects history from being lost. It does not protect history from becoming different. Every simpler answer has now failed. Two more attempts remained.
What if only one machine decides what becomes history, and the rest simply follow? It works — until the authority goes silent. Two remaining machines each reasonably conclude "the authority is gone" and each begin accepting changes.
The architecture recreated the very contradiction it was designed to eliminate: two authorities, two futures.Five machines split into a group of three and a group of two. Each group reasonably believes the other has failed. If both continue accepting writes independently, the contradiction returns. But no group smaller than a majority can ever tell "the rest failed" apart from "we became isolated."
A distributed system cannot build shared truth upon speculation. It must build shared truth upon agreement.Ready: Split five machines and see which group can safely keep accepting writes.
The answer cannot belong to one machine — no single machine possesses complete knowledge. It cannot belong to perfect communication — perfect communication does not exist. The answer must emerge from the collective judgement of the system. A proposed change is not accepted because one machine believes it is correct. It becomes accepted only when enough independent machines reach the same conclusion.
Agreement no longer follows history. History follows agreement. That single inversion changes everything.
Before the shared history is allowed to evolve, the system must first establish that the change has been accepted by the group responsible for protecting that history. Only then does it become part of reality.
Every architectural discovery in this book — controllers, informers, work queues, the source of truth — quietly depended on this without it ever being written down.
Many machines may observe and remember events, but eventually there must exist one authoritative sequence of committed events. Not several competing histories.
Temporary disagreement is unavoidable. Permanent disagreement is not.
Independent branches may appear during failures. The system must converge upon one shared history before future decisions are built.
A machine may observe, speculate, or temporarily lack information, but distributed actions must be grounded in the globally accepted history rather than private memory.
Copying history protects it from loss. Agreement protects it from contradiction. Replication without agreement merely creates multiple copies of uncertainty.
It does not promise instant propagation, that machines never fail, that networks never partition, or that agreement is free. Those are implementation concerns — the responsibility of a consensus mechanism, not of this contract. Kubernetes delegates this to etcd, which implements Raft. The implementation is replaceable. The contract is not.
It is tempting to believe enough networking solves coordination. It does not. Machines can exchange messages endlessly and still disagree about reality — communication moves information, agreement establishes truth.
A search engine, a CDN, or a metrics dashboard can tolerate temporary disagreement in exchange for latency and availability. Infrastructure orchestration, financial transactions, and identity systems cannot.
A system insisting on one shared history may delay decisions until uncertainty resolves, exchange more messages before accepting a change, or refuse to move forward at all. None of this is accidental — it is the price of "everyone builds upon the same past."
Scenario: Three colleagues each keep a notebook of customer orders. One order reaches only two of the three. Another is delayed; everyone keeps recording independently during the delay. Communication is restored — each notebook is internally consistent, and none of them match.
Predict: Which notebook should become the official history? Should they be merged? Should one be chosen — and how?
Observe: Nobody acted maliciously. Nobody's notebook is obviously wrong. The disagreement emerged naturally from independent participants operating without a shared history.
Reflect: The challenge was never storing information. It was deciding which history everyone should trust.
Predict: Production clusters almost always run an odd number of etcd members — three or five, rarely two or four. Why might that be? If one member becomes unreachable, should writes continue? If exactly half become unreachable?
Experiment: kubectl -n kube-system get pods -l component=etcd to see the member count, or use etcdctl member list / endpoint health where direct access is available.
Observe: The commands reveal who exists and who is healthy. They never explain how the cluster decides whether a write is safe enough to become part of the shared history — that mechanism is intentionally hidden here.
Reflect: Somewhere between "one machine decides" and "everyone must agree" lies a rule answering: when is there enough agreement to safely accept new history? We stop one step before discovering that rule.
Consensus guarantees history will not diverge. It does not guarantee that a writer is modifying the latest version of that history. Two controllers can both read real, accepted history, both make individually correct decisions, and both have consensus accept their changes — while one writer's work silently disappears because history moved forward between the read and the write.
Agreement is infrastructure. But agreement only protects what is already accepted. It says nothing about whether the next writer is working from current truth.
The insight that agreement requires more than communication traces back to Fischer, Lynch, and Paterson's 1985 FLP impossibility result: in an asynchronous system, no deterministic consensus protocol can simultaneously guarantee safety, liveness, and tolerance of even one crash failure. Raft and Paxos navigate that boundary by sacrificing liveness under partition — the cluster stops accepting writes rather than allowing histories to diverge. Kubernetes inherits that trade-off through etcd.