Investigation 010 · The Single Leader Problem
Everyone was responsible.
Which meant no one was truly responsible.
Three controllers watch the same Deployment. All three react to the same change. Two of them do wasted, sometimes dangerous, work. Shared state must prevent competing unexpired ownership claims — without assigning ownership forever.
Begin the investigation ↓Prologue
Three controllers, one Deployment
Imagine three replicas of the same controller, all watching the same Deployment. A change arrives. All three observe it. All three attempt to reconcile.
Versioned truth — the contract earned in INV·009 — rejects two of the three writes to the Kubernetes object. Only one succeeds. That much is already solved.
But reconciliation is not only internal writes. Sometimes it means creating a cloud load balancer. Charging a credit card. Sending an email.
Those actions cannot be rejected by a version conflict after the fact. If three replicas each create a load balancer, three load balancers exist. The damage is already done.
First Principles
The coordination challenge
Strip away Kubernetes entirely. Any system that replicates workers for availability, but has some work that cannot safely happen twice, faces the same abstract tension.
Replication demands that many workers stand ready. Exclusive responsibility demands that shared state reject competing unexpired ownership claims. Those two demands pull in opposite directions.
Many replicas remain ready to act at any moment.
Shared state rejects competing unexpired claims to responsibility for work whose effects can’t safely repeat.
The moment a leader can’t act, another takes over immediately.
In the simple design where each follower promotes itself after a local wait, waiting longer reduces premature takeovers but delays recovery; waiting less does the reverse. We have not established whether a different protocol can change that boundary — that remains INV·011’s question.
This is the shape of every design that follows. We are not choosing a clever algorithm. We are choosing which of these three demands we are willing to relax, and when.
We are not choosing a clever algorithm. We are choosing which of these three demands we are willing to relax, and when.
Naive Architecture
The Many Leaders
The simplest architecture requires zero coordination. Every replica watches. Every replica reconciles. No replica asks permission.
all reconcile concurrently
For writes to Kubernetes objects, this mostly survives. Versioned truth rejects the losing writes. Wasted CPU, but no corruption.
For external, non-repeatable effects, it fails outright.
Lab — Three Controllers, One Load Balancer
Fire the same reconciliation event at three uncoordinated controllers and watch what each one does to the outside world.
Waiting for event.
The naive architecture maximizes availability. It sacrifices correctness for anything the API server itself cannot deduplicate. We need an owner.
We need an owner.
The Architecture That Almost Worked
A permanent leader
The obvious fix: permanently designate one controller as leader. Only the leader reconciles. Everyone else stays idle.
standby, do nothing
For the first time, the architecture introduces ownership. The work no longer belongs to every replica — it belongs to exactly one. Every follower now asks “Am I the leader?” instead of “Should I reconcile?”
While every replica follows the ownership rule, external systems receive one request from the designated coordinator. The design appears solved — until an old owner continues after responsibility moves.
The Deployment changes. Both followers observe it. Neither acts — the design says only the leader may act. The cluster is unavailable, not because every controller failed, but because the only controller allowed to work has disappeared.
The first design maximized availability and sacrificed correctness. This one maximizes correctness and sacrifices availability. The real question is no longer who should own the work — it’s when ownership should change.
The real question is no longer who should own the work — it’s when ownership should change.
Breaking Our Design
Failure is not an exceptional event. It is an architectural certainty.
The permanent leader design depends on one assumption: the leader will always remain capable of leading. Reality makes no such promise. We deliberately break the design to expose what it silently assumes.
Episode 1 — The Dead Leader
Controller A crashes. Controllers B and C remain capable, but the permanent-leader architecture contains no transfer rule. If they obey it, both continue waiting.
Discovery: permanent ownership turns one process failure into indefinite unavailability. This episode has not yet tested a follower that promotes itself.
Episode 2 — The Slow Leader
Controller A hasn’t crashed. It’s overloaded — CPU contention, memory swapping, GC pause, saturated disk I/O. It is still running, just slower. To Controller B, today looks exactly like yesterday’s crash: silence.
B concludes A failed and promotes itself. Moments later A catches up and resumes work. Now two leaders exist — not because anything crashed, but because one machine became slower than expected.
Waiting longer changes how often overlap occurs, but it does not create authority to transfer the work. Discovery: a unilateral takeover can overlap with an owner that never stopped acting.
Episode 3 — Silence Cannot Revoke Ownership
Controller B knows only that messages from A have not arrived. A timeout can decide when B stops waiting, but it does not establish why A became silent or whether A stopped acting.
Discovery: silence cannot authorize an ownership transfer by itself. The shared system needs an authoritative rule for when the current ownership claim ends. Whether failure can ever be identified with certainty remains INV·011’s mystery.
Silence cannot authorize an ownership transfer by itself.
Lab — Dead, Slow, or Just Quiet?
A follower only observes silence. Reveal what actually happened behind that silence — and notice that the observation never changes.
Follower B observes: waiting for signal…
Turning Point
Stop letting each follower decide. Put transfer in shared state.
Every failure exposed the same ownership gap: silence did not authorize a follower to revoke the current claim.
The shared system needs one authoritative rule that ends the existing claim before accepting another. We have not yet constructed how that rule ends ownership.
Not local promotion. Shared ownership transfer.
Not local promotion.
Shared ownership transfer.
Episode 4
Ownership, not a house — a rental agreement
The Lease Becomes Inevitable.
Until now, leadership behaved like ownership of a house: acquired once, valid forever until explicitly taken away. That assumption caused every failure above.
Leadership should instead behave like a lease. Acquired, then expiring automatically unless renewed.
If the controller crashes, freezes, or loses access to the shared coordination record, the reason does not need to be diagnosed first. The lease is no longer renewed. It expires on its own. Not because someone proved failure — because the ownership deadline passed.
Once expired, another healthy controller may acquire it. No administrator intervenes. No permanent assignment changes.
Lab — Lease Lifecycle
Healthy Controller A renews automatically. Force a renewal to observe the deadline reset, or crash A to stop renewal and watch the claim expire.
Controller A holds the lease and is renewing normally.
The ownership failures now have a rule. Many cooperative controllers → the shared record recognizes one holder. The permanent owner stops renewing → its claim expires, letting another acquire it. A slow owner → it retains the claim only while renewal succeeds. Followers do not need to diagnose the machine before the claim ends.
The Leadership Contract
Seven responsibilities every correct design must satisfy
No replica should become permanently indispensable.
Shared state recognizes at most one unexpired ownership claim. During transfer, it may recognize none.
Every grant has a limited lifetime. Leadership is a lease, not a title.
Renewal proves only “I still intend to perform this responsibility.”
Ownership expires without requiring anyone to prove failure.
Once expired, another eligible replica may acquire it — no administrator required.
Every replica remains equally capable; only the current owner differs.
Engineering Reflection
Kubernetes did not invent leader election
We deliberately avoided Kubernetes throughout this investigation to discover the underlying architectural principle. Kubernetes expresses temporary ownership using the Lease resource (coordination.k8s.io/v1) together with the client-go leader election library.
Each replica of kube-controller-manager or kube-scheduler attempts to acquire the Lease. The holder becomes active leader and periodically renews it. If renewal stops, the Lease expires and another replica may acquire it.
Critical boundary: the client-go package explicitly states that its leader election does not guarantee fencing. A stale former leader may still be acting, so external effects require their own idempotency, operation identity, or authoritative stale-actor rejection.
Costs Accepted
Renewal overhead. The active leader continuously writes to etcd to renew its lease — background coordination that consumes network and write throughput even when idle.
Leadership gap. When renewal stops, another participant waits for lease expiry before claiming ownership. That gap is the lease’s remaining lifetime plus acquisition time. It reduces premature takeover; it does not prove the former holder stopped acting.
Tuning sensitivity. Lease duration trades recovery speed against false-failover rate. No value is universally correct — this is deliberate policy, outside the architectural contract.
Idle readiness cost. Every standby replica maintains a full informer cache and stays ready, at the cost of memory and watch-event processing, while performing no reconciliation.
These costs are accepted because the alternatives fail harder. A permanently designated leader blocks on failure. Unrestricted concurrent execution duplicates coordination. Temporary ownership with automatic expiry is the least bad coordination option — not a substitute for protecting external effects.
Investigation Exercise
Exercise 1 — The Permanent Leader. Controller A is permanently the leader; B and C are followers. What happens if A crashes? Can B or C safely take over? How would they know A is permanently gone? What if both decide to take over independently?
Exercise 2 — Temporary Ownership. Replace permanent leadership with a lease renewed every few seconds. What happens when renewal succeeds? What happens if it stops? Does anyone need to prove the leader crashed? Compare the failure modes with Exercise 1.
Bridge to INV·011
The lease expired. But why?
We have solved the problem we set out to solve. Many replicas remain available while shared state prevents competing unexpired ownership claims. During transfer, there may be no current holder. Leadership is renewed continuously and expires automatically when renewal stops.
But we quietly relied on one assumption throughout: a lease eventually expires. That tells us leadership was not renewed in time — it tells us nothing about why. Did the leader crash? Pause? Overload? Partition? Experience latency?
We allowed ownership to expire without first diagnosing the holder. Could a smarter protocol reliably distinguish a dead machine from a slow one? Or is uncertainty itself a deep architectural boundary? INV·010 deliberately leaves that unanswered.
That question belongs to its own investigation, its own naive attempt, and its own failure. Not leadership. Not ownership. The nature of failure itself.
Deliberate Simplifications Ledger
| We deliberately postponed | Owned by |
|---|---|
| How the system decides a node has failed rather than merely slowed | INV·011 — Failure Detection & Liveness |
| The optimal lease duration — balancing recovery speed against false failovers | Policy question — implementation detail |
| Stale former holders, in-flight external work, and exactly-once outcomes | The affected domain’s operation-identity, idempotency, or stale-actor rejection contract; the need is formalized in INV·032 |
| The Raft-internal leader election within the etcd consensus layer | INV·008 territory — already derived |
| Clock synchronization between replicas and its effect on lease accuracy | INV·011 — Failure Detection & Liveness |
Sources
Official Documentation: Kubernetes coordination.k8s.io/v1 Lease API reference; kube-controller-manager leader election flags (--leader-elect, --leader-elect-lease-duration, --leader-elect-renew-deadline, --leader-elect-retry-period); Kubernetes Enhancement Proposal — leader election via Lease objects (kubernetes/enhancements/issues/16).
Source Code: k8s.io/client-go/tools/leaderelection; k8s.io/client-go/tools/leaderelection/resourcelock.
Research: The Chubby Lock Service for Loosely-Coupled Distributed Systems — Burrows, OSDI 2006; ZooKeeper: Wait-free coordination for internet-scale systems — Hunt et al., USENIX ATC 2010.
Books: Designing Data-Intensive Applications — Martin Kleppmann (Ch. 8–9); Programming Kubernetes — Hausenblas & Schimanski.
Next: INV·011 — Failure Detection & Liveness