Author's Note

The watch only tells you how one controller learns.

INV-004 replaced polling with a stream: the source of truth speaks, the controller listens, nobody proves nothing happened. This investigation asks what happens when a hundred controllers want to learn about the same change.

What we inherit

Watch, not poll

A controller no longer asks on a schedule. It opens one relationship and reacts when reality moves.

What remains hidden

Observation, shared or private?

Every controller was correct, and every controller was manufacturing the same knowledge alone, as if it were the only one alive.

Not how to observe. We already know that. But what to do when observation is something many components need, and each one is paying the full price alone.
INVESTIGATION 005 / THE SAME FACT, MANY TIMES
Core Control Plane Principles

How Many Times Should One Change Be Observed?

A single Pod is deleted. The ReplicaSet controller receives it. So does the Deployment controller. So does the garbage collector, the endpoint tracker, the autoscaler. One fact, decoded and remembered again and again, by components that all agree on what it means.

If a hundred components need the same fact, shouldn't it be discovered once?
POD
DELETED
ReplicaSet
copy
Deployment
copy
GC
copy
Endpoints
copy
Autoscaler
copy
...
copy
Follow one deletion ↓
Prologue · The Unanswered Question

One deletion, delivered again and again.

Somewhere in the cluster a node reclaims a little memory, and a Pod that used to exist no longer does. It is the smallest possible event. Now watch what happens inside the control plane.

The ReplicaSet controller was watching Pods. It receives the deletion. The Deployment controller was watching Pods too. It receives the same deletion. The garbage collector receives it as well. So does the endpoint tracker. So does the autoscaler. So does every other component with reason to keep an eye on Pods.

Each delivery crosses the boundary between source and observer. Each is decoded from bytes into an object. Each observer updates its own private picture of the cluster. They are all reacting to the same fact, and all doing the same preparatory work to learn it.

The watch already taught us not to ask the same question many times. So why are we still answering the same observation many times?
First Principles

A shared picture, duplicated by effort.

An observer does not just receive events. It reconstructs a current picture of the world from them — the thing controllers actually reason about. Add a second observer, and it wants the very same picture.

The old question

How should a controller observe the cluster?

This assumes observation belongs to each controller, privately reconstructed every time.

The better question

Who should observe, so everyone can use the result?

This treats observation as something the system can own once, and controllers merely consume.

A picture of shared reality is shared by nature. Only the effort of maintaining it has been duplicated.
The Naive Architecture

Everyone Observes Alone

Each controller opens its own watch, reads current state once, follows the stream, and keeps its private picture up to date. This is not a compromise. It is a good design.

Controlleropens own watch
Source of Truthstreams changes
Private Picturereconcile from it
Respects everything we've learned

Fully independent

Each controller depends on no other. If it crashes, it reopens its watch and carries on. Nothing else needs to know or care.

Simple to reason about

One controller, one watch, one picture

No shared components, no coordination. When you debug one controller, you never think about another.

It works. A Pod changes, every controller watching Pods is told, each updates its own picture, each decides independently whether the change matters. On a modest cluster, you would see nothing wrong.

New facts / second200
Decodes / second0
Real Pod objects0
Objects held in memory0

Ready: One fact, decoded and remembered once per observer. Run the numbers.

Why We Should Be Suspicious Anyway

Nothing here is wrong. It is correct at every step. But this book has trained a reflex: when an architecture looks finished and feels effortless, ask what it will cost later — not whether it is correct, but whether its cost is tied to something that will grow.

The Architecture That Almost Worked

Growth feels free.

A handful of controllers, a cluster small enough to fit in memory many times over. Two or three controllers keeping their own copy of a few hundred objects is nothing. At this size, independence is pure profit.

Every addition is clean

Add a controller, it opens its own watch

New resource kinds, autoscaling, endpoint tracking, certificate management — every addition follows the rules perfectly.

The cluster grows too

More Pods, more nodes, more of everything

The shared picture each controller maintains is no longer a few hundred objects. It is tens of thousands — once per controller that cares.

Controllers watching Pods12
Real Pods in cluster30,000
Pod objects held in memory360,000
Objects that exist only for duplication330,000
Eleven out of every twelve of those objects exist only because each observer insisted on keeping its own copy.

Nothing has failed. Every controller is still correct. Every watch still works. The system still converges. But the cost of observation now grows with two things at once: the number of controllers that observe, and the size of the world each observes separately.

In the last investigation, polling made cost grow with curiosity instead of change. Independent watching makes cost grow with the number of observers instead of the amount of reality. Same disease. New organ.

Breaking Our Design

Five costs, one root cause.

We will not break this architecture by finding a bug. There is no bug. We break it the way scale breaks things — by counting.

FAILURE 01

The Same Change, Many Times

Twelve controllers watching Pods means one deletion is decoded twelve times. 200 changes/second becomes 2,400 decodes/second — twelve times the work for exactly the same knowledge.

Decoding scales with reality times the size of the crowd.
FAILURE 02

N Copies of the World

30,000 real Pods become 360,000 Pod objects in memory once twelve controllers each keep their own picture. Every change must then be applied twelve times, not once.

We are treating a shared, read-only fact as if it were private data.
FAILURE 03

A Connection for Every Observer

A watch is a relationship that stays open. An idle cluster with a hundred observers still requires the source to hold a hundred live channels, fed by nothing, for no new facts.

The source becomes a fan-out machine whose load is dominated by how many are listening.
FAILURE 04

The Thundering Herd

When the source restarts or a connection drops, every observer reconnects and re-reads the entire current state at once — a hundred simultaneous rebuilds hitting a source that just came back to life.

Independent components with a common dependency fail independently, but recover together.
FAILURE 05

The Cache That Drifted

Every private picture is late by its own small, independent amount. Twelve observers hold twelve slightly different, momentarily disagreeing versions of one truth.

We are spending twelve times over to get something worse than one honest copy.
Observers reconnecting0
Full re-reads at once0
Shared-informer re-reads1

Ready: Restart the source and watch every independent observer stampede toward the same current state, at the same fragile instant.

The Turning Point

Stop observing separately.

Every fault had one root. Observation was private, but reality was shared.

Source of Truthone watch, not twelve
Shared Observerone remembered picture
Many Controllersread & react, don't observe

Imagine a single component whose only job is to observe one kind of resource. It opens one watch, reads current state once, follows the stream, and keeps one picture — in one place. It does not decide anything. It does not reconcile. It protects no invariant. Its entire purpose is to know, on behalf of everyone who needs to know.

Walk back through the wreckage and watch it collapse: one decode, one memory footprint, one connection, one recovery, one consistent picture. Every symptom traced to duplicated observation. Removing the duplication removes every symptom at once.

Watches on the source12
Decodes / second2,400
Pod objects in memory360,000
Re-reads on recovery12

Independent watches: twelve controllers, twelve of everything — connections, decodes, memory, and simultaneous recovery.

The Architectural Reveal
The architecture earns its nameKubernetes calls it an
Informer.
One watch, one cache, many consumers. When several controllers share the same Informer for the same resource, Kubernetes calls it a shared informer — the arrangement that actually runs inside a control plane.
The Honest Failure

A picture is always a little late.

Every private picture was late by its own small amount before we shared anything. Sharing observation does not remove lateness. It removes the disagreement between copies of the same lateness.

Ready: Compare what each controller believes about the same deleted Pod.

The Observation Fence

Reconciliation is level-triggered, so acting on a slightly stale picture delays convergence — it never corrupts it. An Informer is a cache, and a cache is never the territory. That is why a shared, slightly-behind picture is safe: the safety was earned two investigations ago, not invented here.

The Informer Contract

An Informer observes one kind of resource on behalf of everyone, maintains a single local picture, serves it cheaply, and announces change to everyone who registered interest.

01 / OBSERVE ONCE

One watch per resource kind.

Not one per consumer. The source holds a single channel where it used to hold a dozen.

02 / REMEMBER ONCE

One local picture.

Kept current from that single stream. This is the shared read model everyone consumes.

03 / SERVE READS LOCALLY

Reads never touch the source.

A consumer asking "what exists right now?" reads the local picture. Reads become cheap.

04 / ANNOUNCE & STAY SHARED

Notify, don't decide.

Registered consumers are told when to act. Adding a consumer adds no load to the source of truth.

How It Stays Current

The Informer first reads the complete current state, then follows the change stream and applies each change to its shared picture. If the connection breaks and its previous position cannot be resumed safely, it performs another full read before continuing. The resumable position marker is accepted here as a contract; how that marker is generated and trusted belongs to INV-009.

What an Informer Refuses to Promise

It does not decide anything — deciding remains the controller's job. It does not promise a perfectly current picture — the local picture is a report, assembled with delay. It does not make the source of truth consistent across failing machines. And it does not manage how a consumer absorbs bursts, retries failures, or de-duplicates repeated notifications — that is about to become our problem.

Engineering Reflection

Observe once. Serve many.

Shared, read-only knowledge should be observed once and served to many — not rediscovered independently by every consumer. A database serves clients from one maintained copy. An OS keeps one page cache. A CDN observes an origin once. Different domains, the same architecture.

Architectural honesty

Keep the private watch

A private watch per consumer is fine when only one or two controllers care about a resource, or when consumers live in genuinely separate processes that should stay decoupled.

When the Informer earns its cost

Share the observation

Once many controllers, in the same process, would otherwise each open their own watch and hold their own copy of the same objects.

Costs Accepted

Shared failureIf the one Informer misbehaves, every consumer feels it at once, instead of each controller having its own isolated bad day.
Never-current pictureA cache is a report, not the territory, and must be treated that way at every read.
Deferred mechanismHow the shared observer resumes after losing its place is a contract we accept without inspecting yet.
Deferred consistencyWhether the one watched stream reflects one agreed truth across failing machines is not this investigation's promise.

Investigation Exercise

  1. Prediction
    kube-controller-manager runs dozens of controllers in one binary. Predict how many separate long-lived connections it would hold open just for Pods if each watched independently.
  2. Experiment
    List every controller you already know watches Pods: ReplicaSet, Deployment, Job, DaemonSet, StatefulSet, endpoints, garbage collector, autoscaler.
  3. Observation
    Run the comparison below against a shared Informer serving the same set of consumers.
  4. Reflection
    Explain why the number of controllers watching a resource can grow without the number of watches on that resource growing at all.
PredictionNone
Independent-watch connections0
Shared-informer connections0
Controllers servedHidden

Prediction: Choose one model before running the comparison.

The Next Mystery

Announcements are not work.

Change arrives in bursts. A rollout touches fifty Pods in a second — can the handler keep up?

Reacting can fail. If a handler fails halfway through reacting to a change, does it get another chance?

The same object changes repeatedly. Must the controller reconcile it three times, or only once, for the latest state?

One shared observer serves many consumers while bursts, failed reactions, and repeated announcements leave unfinished work unowned
INV-006

Why Discovery and Execution Must Be Separated

The Informer Contract solves knowing. The next investigation asks how a single controller safely absorbs a stream of announcements without losing work, repeating work, or being overwhelmed.

Deliberate Simplifications Ledger

  • How a consumer absorbs bursts, retries failures, and de-duplicates repeated notificationsINV-006 (Why Discovery and Execution Must Be Separated)
  • How the shared observer resumes correctly after losing its place in the streamINV-009 (Versioned Truth)
  • Why the single watched stream reflects one agreed truth across failing machinesINV-007 (The Source of Truth)
  • Sharing observation across separate processes or machinesNot claimed here — this investigation shares observation within one control plane