Investigation 010 · The Single Leader Problem

Everyone was responsible.
Which meant no one was truly responsible.

Three controllers watch the same Deployment. All three react to the same change. Two of them do wasted, sometimes dangerous, work. Shared state must prevent competing unexpired ownership claims — without assigning ownership forever.

Begin the investigation ↓

Author’s Note

“Leader” means responsibility, not privilege.

In everyday language, a leader sounds special — smarter, more senior, more trusted. Discard that meaning entirely.

In this investigation, a leader is simply whichever replica currently holds responsibility for exclusive work. Any replica must be able to assume that responsibility if the current leader can no longer fulfill it.

Leadership is not a title. It is a temporary assignment.

Prologue

Three controllers, one Deployment

Imagine three replicas of the same controller, all watching the same Deployment. A change arrives. All three observe it. All three attempt to reconcile.

Versioned truth — the contract earned in INV·009 — rejects two of the three writes to the Kubernetes object. Only one succeeds. That much is already solved.

But reconciliation is not only internal writes. Sometimes it means creating a cloud load balancer. Charging a credit card. Sending an email.

Those actions cannot be rejected by a version conflict after the fact. If three replicas each create a load balancer, three load balancers exist. The damage is already done.

The gap. ResourceVersion protects the record of what happened. It cannot protect the world outside that record. Some effects can’t be undone by rejecting a write.

First Principles

The coordination challenge

Strip away Kubernetes entirely. Any system that replicates workers for availability, but has some work that cannot safely happen twice, faces the same abstract tension.

Replication demands that many workers stand ready. Exclusive responsibility demands that shared state reject competing unexpired ownership claims. Those two demands pull in opposite directions.

The impossible goal. Formally, a system cannot simultaneously guarantee all three of the following, perfectly, at all times:
High Availability

Many replicas remain ready to act at any moment.

Exclusive Coordination

Shared state rejects competing unexpired claims to responsibility for work whose effects can’t safely repeat.

Instant Recovery

The moment a leader can’t act, another takes over immediately.

In the simple design where each follower promotes itself after a local wait, waiting longer reduces premature takeovers but delays recovery; waiting less does the reverse. We have not established whether a different protocol can change that boundary — that remains INV·011’s question.

This is the shape of every design that follows. We are not choosing a clever algorithm. We are choosing which of these three demands we are willing to relax, and when.

We are not choosing a clever algorithm. We are choosing which of these three demands we are willing to relax, and when.

Naive Architecture

The Many Leaders

The simplest architecture requires zero coordination. Every replica watches. Every replica reconciles. No replica asks permission.

Deployment Updated
↓
API Server
↓ ↓ ↓
Controller A    Controller B    Controller C
all reconcile concurrently

For writes to Kubernetes objects, this mostly survives. Versioned truth rejects the losing writes. Wasted CPU, but no corruption.

For external, non-repeatable effects, it fails outright.

Lab — Three Controllers, One Load Balancer

Fire the same reconciliation event at three uncoordinated controllers and watch what each one does to the outside world.

Waiting for event.

The naive architecture maximizes availability. It sacrifices correctness for anything the API server itself cannot deduplicate. We need an owner.

We need an owner.

The Architecture That Almost Worked

A permanent leader

The obvious fix: permanently designate one controller as leader. Only the leader reconciles. Everyone else stays idle.

API Server
↓
Leader (A) — reconciles
 
Follower (B)    Follower (C)
standby, do nothing

For the first time, the architecture introduces ownership. The work no longer belongs to every replica — it belongs to exactly one. Every follower now asks “Am I the leader?” instead of “Should I reconcile?”

While every replica follows the ownership rule, external systems receive one request from the designated coordinator. The design appears solved — until an old owner continues after responsibility moves.

The hidden assumption. This design silently assumes the designated leader will always remain available. Distributed systems exist precisely because machines do not live forever.
Leader (A) — CRASHED
↓
Follower B: waiting    Follower C: waiting

The Deployment changes. Both followers observe it. Neither acts — the design says only the leader may act. The cluster is unavailable, not because every controller failed, but because the only controller allowed to work has disappeared.

The first design maximized availability and sacrificed correctness. This one maximizes correctness and sacrifices availability. The real question is no longer who should own the work — it’s when ownership should change.

The real question is no longer who should own the work — it’s when ownership should change.

Breaking Our Design

Failure is not an exceptional event. It is an architectural certainty.

The permanent leader design depends on one assumption: the leader will always remain capable of leading. Reality makes no such promise. We deliberately break the design to expose what it silently assumes.

Episode 1 — The Dead Leader

Controller A crashes. Controllers B and C remain capable, but the permanent-leader architecture contains no transfer rule. If they obey it, both continue waiting.

Discovery: permanent ownership turns one process failure into indefinite unavailability. This episode has not yet tested a follower that promotes itself.

Episode 2 — The Slow Leader

Controller A hasn’t crashed. It’s overloaded — CPU contention, memory swapping, GC pause, saturated disk I/O. It is still running, just slower. To Controller B, today looks exactly like yesterday’s crash: silence.

B concludes A failed and promotes itself. Moments later A catches up and resumes work. Now two leaders exist — not because anything crashed, but because one machine became slower than expected.

Waiting longer changes how often overlap occurs, but it does not create authority to transfer the work. Discovery: a unilateral takeover can overlap with an owner that never stopped acting.

Episode 3 — Silence Cannot Revoke Ownership

Controller B knows only that messages from A have not arrived. A timeout can decide when B stops waiting, but it does not establish why A became silent or whether A stopped acting.

Discovery: silence cannot authorize an ownership transfer by itself. The shared system needs an authoritative rule for when the current ownership claim ends. Whether failure can ever be identified with certainty remains INV·011’s mystery.

Silence cannot authorize an ownership transfer by itself.
Compact review

Lab — Dead, Slow, or Just Quiet?

A follower only observes silence. Reveal what actually happened behind that silence — and notice that the observation never changes.

Follower B observes: waiting for signal…

Turning Point

Stop letting each follower decide. Put transfer in shared state.

Every failure exposed the same ownership gap: silence did not authorize a follower to revoke the current claim.

The shared system needs one authoritative rule that ends the existing claim before accepting another. We have not yet constructed how that rule ends ownership.

Not local promotion. Shared ownership transfer.

Not local promotion.
Shared ownership transfer.

Episode 4

Ownership, not a house — a rental agreement

The Lease Becomes Inevitable.

Until now, leadership behaved like ownership of a house: acquired once, valid forever until explicitly taken away. That assumption caused every failure above.

Leadership should instead behave like a lease. Acquired, then expiring automatically unless renewed.

Leader granted lease — valid until 10:30:00
↓ renew before deadline
Valid until 10:30:15 — leadership continues

If the controller crashes, freezes, or loses access to the shared coordination record, the reason does not need to be diagnosed first. The lease is no longer renewed. It expires on its own. Not because someone proved failure — because the ownership deadline passed.

Once expired, another healthy controller may acquire it. No administrator intervenes. No permanent assignment changes.

A cooperative controller may initiate lease-guarded work only while it holds the current ownership claim. The rule coordinates responsibility without treating silence as proof of failure.

Lab — Lease Lifecycle

Healthy Controller A renews automatically. Force a renewal to observe the deadline reset, or crash A to stop renewal and watch the claim expire.

Holder: Controller A15.0s remaining

Controller A holds the lease and is renewing normally.

The ownership failures now have a rule. Many cooperative controllers → the shared record recognizes one holder. The permanent owner stops renewing → its claim expires, letting another acquire it. A slow owner → it retains the claim only while renewal succeeds. Followers do not need to diagnose the machine before the claim ends.

Lease ownership is not fencing. The record cannot force a paused or partitioned former holder to stop, retract in-flight work, or make an external effect exactly once. Stronger protection must be enforced at the affected resource boundary.

The Leadership Contract

Seven responsibilities every correct design must satisfy

1. Many replicas remain ready

No replica should become permanently indispensable.

2. No competing claims

Shared state recognizes at most one unexpired ownership claim. During transfer, it may recognize none.

3. Leadership is temporary

Every grant has a limited lifetime. Leadership is a lease, not a title.

4. Continuous renewal

Renewal proves only “I still intend to perform this responsibility.”

5. Automatic expiration

Ownership expires without requiring anyone to prove failure.

6. Transferable responsibility

Once expired, another eligible replica may acquire it — no administrator required.

7. Leadership ≠ capability

Every replica remains equally capable; only the current owner differs.

The timeless principle. Distributed systems achieve high availability by replicating workers, then coordinate exclusive responsibility by preventing temporary authoritative claims from overlapping in shared state. That record coordinates cooperative participants; it does not physically exclude stale actors.

Engineering Reflection

Kubernetes did not invent leader election

We deliberately avoided Kubernetes throughout this investigation to discover the underlying architectural principle. Kubernetes expresses temporary ownership using the Lease resource (coordination.k8s.io/v1) together with the client-go leader election library.

Each replica of kube-controller-manager or kube-scheduler attempts to acquire the Lease. The holder becomes active leader and periodically renews it. If renewal stops, the Lease expires and another replica may acquire it.

Critical boundary: the client-go package explicitly states that its leader election does not guarantee fencing. A stale former leader may still be acting, so external effects require their own idempotency, operation identity, or authoritative stale-actor rejection.

Costs Accepted

Renewal overhead. The active leader continuously writes to etcd to renew its lease — background coordination that consumes network and write throughput even when idle.

Leadership gap. When renewal stops, another participant waits for lease expiry before claiming ownership. That gap is the lease’s remaining lifetime plus acquisition time. It reduces premature takeover; it does not prove the former holder stopped acting.

Tuning sensitivity. Lease duration trades recovery speed against false-failover rate. No value is universally correct — this is deliberate policy, outside the architectural contract.

Idle readiness cost. Every standby replica maintains a full informer cache and stays ready, at the cost of memory and watch-event processing, while performing no reconciliation.

These costs are accepted because the alternatives fail harder. A permanently designated leader blocks on failure. Unrestricted concurrent execution duplicates coordination. Temporary ownership with automatic expiry is the least bad coordination option — not a substitute for protecting external effects.

Investigation Exercise

Exercise 1 — The Permanent Leader. Controller A is permanently the leader; B and C are followers. What happens if A crashes? Can B or C safely take over? How would they know A is permanently gone? What if both decide to take over independently?

Exercise 2 — Temporary Ownership. Replace permanent leadership with a lease renewed every few seconds. What happens when renewal succeeds? What happens if it stops? Does anyone need to prove the leader crashed? Compare the failure modes with Exercise 1.

Bridge to INV·011

The lease expired. But why?

We have solved the problem we set out to solve. Many replicas remain available while shared state prevents competing unexpired ownership claims. During transfer, there may be no current holder. Leadership is renewed continuously and expires automatically when renewal stops.

But we quietly relied on one assumption throughout: a lease eventually expires. That tells us leadership was not renewed in time — it tells us nothing about why. Did the leader crash? Pause? Overload? Partition? Experience latency?

Leadership and failure are related, not identical. Leadership answers “who is currently allowed to perform exclusive work?” Failure detection answers “can this machine still be trusted to participate?”

We allowed ownership to expire without first diagnosing the holder. Could a smarter protocol reliably distinguish a dead machine from a slow one? Or is uncertainty itself a deep architectural boundary? INV·010 deliberately leaves that unanswered.

That question belongs to its own investigation, its own naive attempt, and its own failure. Not leadership. Not ownership. The nature of failure itself.

Shared state recognizes at most one current claim to exclusive responsibility, but missing renewal cannot distinguish a crash, pause, partition, or delay

Deliberate Simplifications Ledger

We deliberately postponedOwned by
How the system decides a node has failed rather than merely slowedINV·011 — Failure Detection & Liveness
The optimal lease duration — balancing recovery speed against false failoversPolicy question — implementation detail
Stale former holders, in-flight external work, and exactly-once outcomesThe affected domain’s operation-identity, idempotency, or stale-actor rejection contract; the need is formalized in INV·032
The Raft-internal leader election within the etcd consensus layerINV·008 territory — already derived
Clock synchronization between replicas and its effect on lease accuracyINV·011 — Failure Detection & Liveness

Sources

Official Documentation: Kubernetes coordination.k8s.io/v1 Lease API reference; kube-controller-manager leader election flags (--leader-elect, --leader-elect-lease-duration, --leader-elect-renew-deadline, --leader-elect-retry-period); Kubernetes Enhancement Proposal — leader election via Lease objects (kubernetes/enhancements/issues/16).

Source Code: k8s.io/client-go/tools/leaderelection; k8s.io/client-go/tools/leaderelection/resourcelock.

Research: The Chubby Lock Service for Loosely-Coupled Distributed Systems — Burrows, OSDI 2006; ZooKeeper: Wait-free coordination for internet-scale systems — Hunt et al., USENIX ATC 2010.

Books: Designing Data-Intensive Applications — Martin Kleppmann (Ch. 8–9); Programming Kubernetes — Hausenblas & Schimanski.

Next: INV·011 — Failure Detection & Liveness