One shared Informer
Every controller now observes the cluster through the same cache instead of watching independently.
A dispatcher answers an emergency call, then personally drives to the incident before answering the phone again. The moment a storm arrives, calls pile up faster than any one person can resolve them. Kubernetes controllers face the same trap.
Discovering that work exists and completing it are not the same responsibility.
A city rule says the dispatcher who answers an emergency call must personally drive to the incident, solve it, and return before answering the next call. One emergency at a time, the system works fine. Then a storm arrives.
Hundreds of emergencies begin arriving every minute. While the dispatcher is extinguishing one fire, dozens of new calls continue arriving. Some callers hang up before anyone answers. Others call again because no one responded. The dispatch center slowly loses track of what still needs attention.
None of the failures happen because firefighters are slow. The city confused two completely different responsibilities: discovering that work exists, and performing the work.
The moment work is discovered is rarely the right moment to perform it.
Discovery answers "has something changed?" and should happen immediately. Execution answers "can I safely finish the work?" and consumes time — sometimes milliseconds, sometimes much longer.
Its value comes from noticing reality as quickly as possible. It does not ask whether execution has finished.
It may succeed, fail, or need to be attempted again. It cannot be rushed just because discovery moved on.
If discovering a change takes one millisecond and completing the work takes one second, and changes arrive every millisecond: in one second, 1,000 changes are discovered and 1 piece of work is completed.
Nothing is broken. No hardware failed. The mathematics alone guarantees that execution cannot keep pace with observation. Someone must answer a surprisingly important question: who is responsible for remembering that this work still exists?
An object changes. The moment the system notices, it immediately begins the work required to respond. Nothing waits. Nothing is remembered. Nothing is scheduled for later.
The observer discovers the work and starts the work. No additional machinery, no coordination, no question of ownership.
It quietly assumes discovering work and performing work happen at nearly the same speed. Reality does not promise this.
Ready: Configure discovery and execution speed, then run one second of a busy cluster.
If discovering work and completing work happen at different speeds, who owns the work in between? Not the Informer — it has already moved on. Not a worker — it may be busy with something else. Yet the work has not disappeared.
Informers already solved observation. The obvious next step: the Informer should immediately tell the controller to reconcile. Nothing waits. Nothing is buffered.
The work completes. The controller waits for the next change. Everything remains simple — every responsibility feels well defined.
Reconciling may require reading other resources, comparing state, creating or deleting objects, and waiting for responses. Each reconciliation consumes time.
The Informer has discovered another change. The previous reconciliation has not yet finished. Should the handler start another reconciliation? Should it wait? Should it ignore the new observation? None of these answers feel entirely satisfactory. The design has not failed. Not yet.
If observation continues while execution is still busy, who remembers that unfinished work still exists?
We are not looking for implementation bugs. We are looking for architectural assumptions — each reasonable on a quiet afternoon, each increasingly difficult to defend as the system grows.
An Informer discovering 1,000 changes/second against a worker completing 10 reconciliations/second leaves 990 discovered changes waiting, every single second.
Every discovered change can be processed immediately — false.A Deployment needs three new Pods. The first is created; before the second, the API server becomes unavailable. The handler already returned — the event is over, but the work is not.
Every reconciliation succeeds — false.A Deployment changes four times before reconciliation starts. Only the latest desired state matters — the controller needs to converge once, not visit every historical value.
Every observation represents unique work — false.2,000 changes/second discovered; even ten workers complete only about 200 reconciliations/second. The gap does not disappear — it accumulates, with nowhere to live.
Execution naturally keeps pace with observation — false.The Informer's responsibility ends at discovery. The handler's callback already returned. Neither owns the work while it waits — like a relay baton that belongs to no one between runners.
A responsibility without an explicit owner is a dangerous place to be.None of these failures were caused by the Informer or the controller — both did exactly what they were designed to do. The architecture itself never built anything responsible for owning work in between.
The Direct Reaction Machine is not inefficient. It is incomplete.Ready: Trigger three rapid notifications about the same object and compare both architectures.
Not a faster observer. Not a smarter controller. A component whose only job is to own unfinished work until reconciliation succeeds.
The moment an observation suggests that an object may require reconciliation, the system records that responsibility somewhere safe. From that point on, the responsibility no longer belongs to the Informer. When a worker becomes available, it offers the next piece of work. If reconciliation succeeds, it forgets the work. If it fails, it keeps remembering.
Direct Reaction: discovery and execution share one fate — whatever happens to one happens to the other.
The queue is not important because it stores items. It is important because it owns unfinished work.
Observation discovers work. The Work Queue remembers it, until reconciliation succeeds.
The queue does not decide whether reconciliation is necessary. It accepts ownership of the possibility that work exists.
Workers may become busy, fail, or need multiple attempts. The work continues to exist until reconciliation succeeds.
Observation does not wait for workers. Workers do not interrupt observation. Each proceeds independently.
Workers never search the cluster. Whenever a worker is available, the queue provides the next object requiring attention.
Only after desired and actual state agree may the queue forget the work — no sooner, no later.
The queue records the identity of work that needs attention, not the count of times it was announced. A channel preserves every message; a Work Queue preserves the responsibility to check an object once.
It does not observe the cluster — that belongs to the Informer. It does not determine the desired state — that belongs to reconciliation. It does not modify cluster resources — that belongs to the controller. It exists for one reason: to own unfinished work until reconciliation succeeds. Nothing more. Nothing less.
An OS separates a hardware interrupt from process scheduling. A network stack separates packet arrival from packet processing. A CPU separates instruction fetch from execution. A message broker separates production from consumption. Kubernetes applies the same timeless pattern to reconciliation.
Fine when change is rare and reconciliation is fast and reliable. A Work Queue earns its cost once change arrives in bursts, reconciliation can fail, or the same object can change again before it has been handled once.
Once bursts, failures, and repeated notifications become ordinary, correctness can no longer depend on perfect timing.
Prediction: Choose one model before running the comparison.
If a controller loses its watch connection, how does it resume from the correct point?
If several controllers observe the same object, how do they all agree on which version of reality they are seeing?
If an object changes thousands of times, how does the system distinguish new information from information already processed?