Author's Note

Six investigations asked you to trust a single word.

The. The cluster. The current replica count. The actual state a reconciler observes. Every earlier investigation quietly assumed there was exactly one of it. This investigation finally asks why that assumption was ever safe to make.

What we inherit

A shared observation pipeline

Informers, Work Queues, and reconciliation all assumed a single, unambiguous "actual state" waiting to be read.

What remains hidden

Where that single picture comes from

Something works hard, in a place none of the previous investigations looked, to make many machines behave as if they were one.

Can more than one machine ever agree on anything, given that any one of them might fail at any moment?
INVESTIGATION 007 / ONE CONSISTENT STORE, OR THE ILLUSION OF ONE
Core Control Plane Principles

Where Does "The Cluster" Actually Live?

A ReplicaSet notices a missing Pod. An Informer keeps a cache dozens of controllers trust. A Work Queue remembers a failed reconciliation. Every one of them assumed there was exactly one cluster to read from, and that two controllers asking the same question get the same answer.

What happens the moment more than one machine is responsible for holding that single, shared picture of reality?
API ServerONE MEMORY
API ServerDOES NOT EXIST YET
The 2:14 a.m. page ↓
Prologue · The 2:14 A.M. Page

Where was the desired state, while the server was down?

"The API Server is returning errors." Every kubectl get hangs. Every controller's watch has gone quiet. The engineer restarts the process. Everything resumes, as if nothing happened.

Then a colleague asks the question that ruins the rest of the night. "Where was the desired state, while the server was down?" — "...On the server." — "The one that was down?" — "...Yes." — "So for four minutes, there was no cluster. There was just an outage, and a very patient set of controllers waiting for someone to pick up the phone again."

"We should run two API Servers," someone finally says. "For availability." It sounds like the obvious fix. It is about to become the beginning of a much longer conversation.
First Principles

Durability and agreement are not the same thing.

A controller's reconciliation only means something if whoever reads "actual state," and whenever they read it, is reading the same fact. Call this agreement — not a formal consensus algorithm yet, just two readers getting the same answer.

Agreement

One memory location has it for free.

There is only one copy, so there is nothing to disagree with.

Durability

One memory location has none.

It can disappear in exactly one event: one crashed process, one severed cable.

A single machine gives you agreement for free, and durability not at all. Copying that machine's memory onto several machines is the obvious answer to durability. It is not obviously an answer to agreement.
The Naive Architecture

The One True Copy

One API Server. One process, on one machine, holding the entire desired and actual state of the cluster in memory. Every controller from every earlier investigation plugs into this without modification.

ControllersA, B, C...
One API Serversingle process
One Memorysingle copy
Why this works, completely, for a while

Nothing to disagree with

If Controller A and B ask the same question one millisecond apart, they get the same answer, because they are asking the same memory, not two memories that are supposed to match.

The cost hiding inside the simplicity

Everything stops

Not "everything degrades." Every controller is, at that instant, unable to observe reality and unable to correct it, because the one place reality lived is gone.

API ServerRunning
BackupDoes not exist
Controllers observingYes
Controllers reconcilingYes
Desired state reachableYes

Ready: This is the 2:14 a.m. page, reproduced on demand. Kill the one machine holding the truth.

The Architecture That Almost Worked

Two API Servers.

Identical software, identical configuration, each holding its own copy of cluster state. A load balancer sends each request to whichever one is free.

Controllersevery request
Load Balancerpicks one
Server 1 or 2own memory each

The team schedules a chaos drill. They kill Server 1 mid-afternoon, on purpose, in front of an audience. kubectl get pods doesn't even hiccup. "We just made the API Server highly available," someone says, "and it cost us almost nothing."

Then a quiet voice near the back asks: "During the drill, was anyone writing to the cluster?" — "...No. We drained traffic first." — "So we tested what happens when nobody writes anything while a server is down."

The chaos drill proved a crash no longer stops the cluster. It did not prove that the two copies agree. Two servers never written to while one is down will look identical — that is silence mistaken for agreement.
Breaking Our Design

Five failures, no bad actors.

The naive architecture survives the failure that started this investigation: one server going down no longer stops the cluster. It does not survive anything else.

FAILURE 01

The Write That Arrived Late

A controller scales a Deployment to 5 replicas on Server 1. Milliseconds later, a dashboard controller reads Server 2 and gets 3. Nobody made a mistake — Server 2 simply hasn't heard yet.

For that window, there is no single truth — only two, briefly disagreeing.
FAILURE 02

The Read That Contradicted the Write

With 400ms of extra latency, two correct reconciliation loops each see a different replica count and each create three new Pods. Together they create eight Pods where five were desired.

Reconciliation was never designed to protect you from disagreeing about what "actual state" currently is.
FAILURE 03

Both Servers Say Yes

Two clients scale the same Deployment to 5 and to 2 at nearly the same instant, one request landing on each server. Each server accepts its write honestly, by its own rules.

A system where every member can unilaterally decide something is true has no way to prevent two members deciding two different things.
FAILURE 04

The Partition

A network split leaves each server healthy, reachable, and correct — from where it's sitting. Each keeps accepting writes from the controllers that can still reach it. Two legitimate histories diverge.

A partitioned server cannot distinguish "I am the only one still working" from "I am the one who got isolated."
FAILURE 05

More Copies Is Not More Truth

We solved a single point of failure by adding a second server, and quietly created a second problem we never had: two servers can disagree, and neither is wrong by its own local rules.

Redundancy is not an agreement protocol. It is just more opinions.
Server 1replicas = 3
Server 2replicas = 3

Ready: Write to one server, then read the other before replication catches up.

Server 1replicas = 3
Server 2replicas = 3

Ready: Two clients each scale the same Deployment differently, one request per server. See which one is "wrong."

The Turning Point

One History, Not Many Copies

We do not need more copies. We need exactly one thing: many machines producing one history that all of them can trust.

Fallible Machinesslow, disconnected, wrong
Agreement Firstbefore a write succeeds
One Authoritative Historydurable & agreed

A distributed control plane does not need multiple copies of the truth. It needs one authoritative history of accepted changes — durable enough to survive any single machine's failure, and agreed upon by more than any single machine's opinion.

This sentence does not say how many machines must agree, or what happens if exactly half are reachable. Those are not omissions. They are the next mystery, and it deserves an investigation that isn't rushed.

Ready: Compare what "history" means with independent copies versus one agreed record.

The Architectural Reveal
Not a machine. A history.One authoritative
history.
Not because any one machine is trustworthy alone — but because enough of them agree. How that agreement is actually reached is the next investigation.
The Source of Truth Contract

One history, not one machine.

Every controller, Informer, and Work Queue from the previous six investigations has been quietly relying on this without it ever being written down.

01

One history, not one machine

The cluster's truth is not "whatever this process remembers." It is one authoritative sequence of accepted changes, which may be stored across many machines.

02

Durability without disagreement

Surviving the loss of a machine must never come at the cost of two machines believing different, incompatible things about the same history.

03

Agreement precedes durability

A write is not "safe" merely because it was copied somewhere. It is safe only after the system has established one accepted outcome before any client is told it succeeded.

What this contract refuses to promise

It does not say how many machines participate, how they establish one accepted outcome, or what they do while communication is interrupted. It does not promise instant agreement. And it does not eliminate reconciliation: controllers still observe a delayed report, but now that report belongs to one authoritative history rather than a private, competing copy.

Engineering Reflection

Durability and agreement are orthogonal, not opposites.

You can have durability without agreement — many copies that disagree. You can have agreement without durability — one machine that agrees with itself perfectly, until it disappears. The architecture we're heading toward needs both, simultaneously, on purpose.

Architectural Honesty

A single server is sometimes the right answer

For a low-stakes internal tool where an outage is an inconvenience, not an incident, one machine with backups may be entirely appropriate. The failure mode is honest: it is down, or it is up.

Costs Accepted

Agreement costs latency, on purpose

Every write that must be agreed upon before it is durable is slower than a write that is merely stored. That latency is not a bug being tolerated — it is the price of the guarantee.

Investigation Exercise

Predict: If the API Server's backing store becomes unreachable, what will kubectl get pods do — hang, error immediately, or return stale data?

Experiment: On a real cluster, block access from the API Server to its storage layer (or stop it), then run kubectl get pods and watch existing controllers.

Observe: Does the API Server serve cached reads, refuse writes, or refuse everything? Do already-running controllers keep reconciling from their last known state?

Reflect: Which of the API Server's behaviors were protecting durability, and which were protecting agreement? Were they ever the same protection?

Bridge to Investigation 008

The Consensus Problem

We now know what we need: one authoritative history, durable across machine failure, agreed upon rather than merely copied. We do not yet know how several independent, fallible machines actually reach that agreement — especially when some of them cannot hear each other at all.

How agreement is reached, what participation it requires, and how disconnected machines rejoin the accepted history remain unanswered.

Next: Investigation 008

The Consensus Problem

How do multiple machines, any of which might crash or be partitioned, agree on a single sequence of events — without ever being told to trust each other?

Several fallible machines must preserve one authoritative history, leaving unresolved how they can agree when messages are delayed or lost

Intellectual Lineage

The problem of agreement among fallible, disconnected machines predates Kubernetes. INV-008 will derive that problem on its own terms before naming any particular mechanism.

Deliberate Simplifications Ledger

  • How fallible machines actually reach agreement — deferred to INV-008, The Consensus Problem.
  • What Kubernetes specifically uses to satisfy the one-history contract — deferred to the end of INV-008, after the mechanism is earned.
  • How independent writers avoid silently overwriting one another once history is agreed — deferred to INV-009, Versioned Truth.
  • What stops two copies of the same controller from both acting on the agreed history — deferred to INV-010, The Single Leader Problem.
  • How the system decides that a machine failed rather than merely went quiet — deferred to INV-011, Failure Detection.

Sources

  • Kubernetes documentation — API Server and cluster architecture overview.
  • Kubernetes kube-apiserver storage interface source (k8s.io/apiserver/pkg/storage).
  • Jeffrey Dean & Sanjay Ghemawat et al., "Large-scale cluster management at Google with Borg" (EuroSys 2015).
  • Martin Kleppmann, Designing Data-Intensive Applications, Chapters 5 & 9 (Replication, Consistency and Consensus).