Author's Note

Discovering work is not the same as owning it.

INV-005 made observation shared infrastructure: one Informer, one cache, many readers. This investigation asks what should happen the instant that shared Informer notices a change — and why reacting immediately is a trap.

What we inherit

One shared Informer

Every controller now observes the cluster through the same cache instead of watching independently.

What remains hidden

Who remembers unfinished work?

The Informer discovers a change and moves on. Reconciliation takes time. Between those two moments, something must own the work.

The discovery of work should not depend on the completion of work.
INVESTIGATION 006 / THE DISPATCHER WHO TRIED TO DO EVERYTHING
Core Control Plane Principles

Once Work Is Discovered, Who Remembers It?

A dispatcher answers an emergency call, then personally drives to the incident before answering the phone again. The moment a storm arrives, calls pile up faster than any one person can resolve them. Kubernetes controllers face the same trap.

Discovering that work exists and completing it are not the same responsibility.
DISCOVERimmediate
? ? ?who remembers
EXECUTEtakes time
Meet the dispatcher ↓
Prologue · The Dispatcher Who Tried to Do Everything

Every second on the fire is a second not spent noticing.

A city rule says the dispatcher who answers an emergency call must personally drive to the incident, solve it, and return before answering the next call. One emergency at a time, the system works fine. Then a storm arrives.

Hundreds of emergencies begin arriving every minute. While the dispatcher is extinguishing one fire, dozens of new calls continue arriving. Some callers hang up before anyone answers. Others call again because no one responded. The dispatch center slowly loses track of what still needs attention.

None of the failures happen because firefighters are slow. The city confused two completely different responsibilities: discovering that work exists, and performing the work.

The moment work is discovered is rarely the right moment to perform it.
First Principles

Discovery and execution obey different laws.

Discovery answers "has something changed?" and should happen immediately. Execution answers "can I safely finish the work?" and consumes time — sometimes milliseconds, sometimes much longer.

Awareness

Discovery is a moment.

Its value comes from noticing reality as quickly as possible. It does not ask whether execution has finished.

Progress

Execution takes time.

It may succeed, fail, or need to be attempted again. It cannot be rushed just because discovery moved on.

If discovering a change takes one millisecond and completing the work takes one second, and changes arrive every millisecond: in one second, 1,000 changes are discovered and 1 piece of work is completed.

Nothing is broken. No hardware failed. The mathematics alone guarantees that execution cannot keep pace with observation. Someone must answer a surprisingly important question: who is responsible for remembering that this work still exists?

The Naive Architecture

The Direct Reaction Machine

An object changes. The moment the system notices, it immediately begins the work required to respond. Nothing waits. Nothing is remembered. Nothing is scheduled for later.

Something Changesobserved instantly
Discover the Changeinformer notices
React Immediatelyreconcile now
Deeply appealing

Minimizes responsibility

The observer discovers the work and starts the work. No additional machinery, no coordination, no question of ownership.

Works beautifully, until...

One hidden assumption

It quietly assumes discovering work and performing work happen at nearly the same speed. Reality does not promise this.

Discovered / second0
Completed / second0
Backlog after 1 second0
Who remembers the backlog?—

Ready: Configure discovery and execution speed, then run one second of a busy cluster.

The Question That Shapes the Architecture

If discovering work and completing work happen at different speeds, who owns the work in between? Not the Informer — it has already moved on. Not a worker — it may be busy with something else. Yet the work has not disappeared.

The Architecture That Almost Worked

Applied to Kubernetes, it looks complete.

Informers already solved observation. The obvious next step: the Informer should immediately tell the controller to reconcile. Nothing waits. Nothing is buffered.

Informerobserves cluster
Event Handlerreacts instantly
Reconcile()restores state
For a quiet cluster

One change, one reconciliation

The work completes. The controller waits for the next change. Everything remains simple — every responsibility feels well defined.

The inherited assumption

Discovery and completion move together

Reconciling may require reading other resources, comparing state, creating or deleting objects, and waiting for responses. Each reconciliation consumes time.

The Informer has discovered another change. The previous reconciliation has not yet finished. Should the handler start another reconciliation? Should it wait? Should it ignore the new observation? None of these answers feel entirely satisfactory. The design has not failed. Not yet.

If observation continues while execution is still busy, who remembers that unfinished work still exists?
Breaking Our Design

Six pressures, one missing owner.

We are not looking for implementation bugs. We are looking for architectural assumptions — each reasonable on a quiet afternoon, each increasingly difficult to defend as the system grows.

PRESSURE 01

The Burst That Overwhelmed Us

An Informer discovering 1,000 changes/second against a worker completing 10 reconciliations/second leaves 990 discovered changes waiting, every single second.

Every discovered change can be processed immediately — false.
PRESSURE 02

When Reconciliation Fails

A Deployment needs three new Pods. The first is created; before the second, the API server becomes unavailable. The handler already returned — the event is over, but the work is not.

Every reconciliation succeeds — false.
PRESSURE 03

The Same Object Changed Again

A Deployment changes four times before reconciliation starts. Only the latest desired state matters — the controller needs to converge once, not visit every historical value.

Every observation represents unique work — false.
PRESSURE 04

Fast Events, Slow Workers

2,000 changes/second discovered; even ten workers complete only about 200 reconciliations/second. The gap does not disappear — it accumulates, with nowhere to live.

Execution naturally keeps pace with observation — false.
PRESSURE 05

The Forgotten Change

The Informer's responsibility ends at discovery. The handler's callback already returned. Neither owns the work while it waits — like a relay baton that belongs to no one between runners.

A responsibility without an explicit owner is a dangerous place to be.
PRESSURE 06

Breaking Direct Reaction

None of these failures were caused by the Informer or the controller — both did exactly what they were designed to do. The architecture itself never built anything responsible for owning work in between.

The Direct Reaction Machine is not inefficient. It is incomplete.
Notifications observed0
Direct-reaction reconciles0
Work-queue reconciles0
Correct number needed1

    Ready: Trigger three rapid notifications about the same object and compare both architectures.

    The Turning Point

    Invent the missing responsibility.

    Not a faster observer. Not a smarter controller. A component whose only job is to own unfinished work until reconciliation succeeds.

    Informerdiscovers work
    Work Queueremembers work
    Workerexecutes work

    The moment an observation suggests that an object may require reconciliation, the system records that responsibility somewhere safe. From that point on, the responsibility no longer belongs to the Informer. When a worker becomes available, it offers the next piece of work. If reconciliation succeeds, it forgets the work. If it fails, it keeps remembering.

    Owns unfinished workNo one
    Repeated notificationsReconciled 3x
    On reconciliation failureWork vanishes
    Observation blocked?Effectively yes

    Direct Reaction: discovery and execution share one fate — whatever happens to one happens to the other.

    The Architectural Reveal
    The architecture earns its nameKubernetes calls it a
    Work Queue.
    The queue is not important because it stores items. It is important because it owns unfinished work.
    The Work Queue Contract

    Observation discovers work. The Work Queue remembers it, until reconciliation succeeds.

    01 / ACCEPT

    Accept discovered work.

    The queue does not decide whether reconciliation is necessary. It accepts ownership of the possibility that work exists.

    02 / PRESERVE

    Preserve ownership until success.

    Workers may become busy, fail, or need multiple attempts. The work continues to exist until reconciliation succeeds.

    03 / DECOUPLE

    Decouple observation from execution.

    Observation does not wait for workers. Workers do not interrupt observation. Each proceeds independently.

    04 / PRESENT

    Present work to workers.

    Workers never search the cluster. Whenever a worker is available, the queue provides the next object requiring attention.

    05 / FORGET

    Forget completed work.

    Only after desired and actual state agree may the queue forget the work — no sooner, no later.

    06 / COLLAPSE

    Collapse repeated interest.

    The queue records the identity of work that needs attention, not the count of times it was announced. A channel preserves every message; a Work Queue preserves the responsibility to check an object once.

    What the Work Queue Does Not Do

    It does not observe the cluster — that belongs to the Informer. It does not determine the desired state — that belongs to reconciliation. It does not modify cluster resources — that belongs to the controller. It exists for one reason: to own unfinished work until reconciliation succeeds. Nothing more. Nothing less.

    Engineering Reflection

    Separate discovering, remembering, executing.

    An OS separates a hardware interrupt from process scheduling. A network stack separates packet arrival from packet processing. A CPU separates instruction fetch from execution. A message broker separates production from consumption. Kubernetes applies the same timeless pattern to reconciliation.

    Architectural honesty

    React directly in the handler

    Fine when change is rare and reconciliation is fast and reliable. A Work Queue earns its cost once change arrives in bursts, reconciliation can fail, or the same object can change again before it has been handled once.

    When the queue earns its cost

    Own unfinished work explicitly

    Once bursts, failures, and repeated notifications become ordinary, correctness can no longer depend on perfect timing.

    Costs Accepted

    A monitored gapBetween "the controller knows about the change" and "the controller has acted on it" — queue depth, retries, and age of the oldest item now need watching.
    Extra machineryA dedicated component exists solely to remember work, adding a moving part that a quiet system would not need.
    Deferred retry policyHow aggressively a failing item is retried is a separate economic question this investigation does not answer.
    Deferred agreementWhether the single stream every queue is fed from reflects one agreed truth across failing machines is not this investigation's promise.

    Investigation Exercise

    1. Prediction
      If the same Pod changes three times in one second, and a controller reacts immediately in its handler with no queue in between, how many times does it reconcile that Pod?
    2. Experiment
      Trace the three notifications through a direct-reaction design, then through a queue keyed by object identity, where adding an already-queued key is a no-op.
    3. Observation
      Compare reconciliation counts between the two designs for the same three notifications.
    4. Reflection
      Explain why the queue didn't make reconciliation faster — it made repeated notifications stop being repeated work.
    PredictionNone
    Direct-reaction reconciles0
    Work-queue reconciles0
    State reconciled againstHidden

    Prediction: Choose one model before running the comparison.

    The Next Mystery

    Every component trusted the same history.

    If a controller loses its watch connection, how does it resume from the correct point?

    If several controllers observe the same object, how do they all agree on which version of reality they are seeing?

    If an object changes thousands of times, how does the system distinguish new information from information already processed?

    Discovery, work ownership, and execution are separated while every component still depends on one trusted history
    INV-007

    The Source of Truth

    Every Informer, Work Queue, and worker in this investigation quietly trusted one consistent history of the cluster. The next investigation asks why that trust is ever justified.

    Deliberate Simplifications Ledger

    • How aggressively a repeatedly failing item should be retriedINV-039 (The Retry Economy Problem)
    • How many workers should drain the queue, and how that number is chosenINV-039 (The Retry Economy Problem)
    • How the resumable position marker a controller resumes from is generated and trustedINV-009 (Versioned Truth)
    • Why the single stream every Work Queue is fed from reflects one agreed truth across failing machinesINV-007 / INV-008
    • What stops two controllers from reconciling the same object at the same instant and clobbering each other's resultINV-009 (Versioned Truth)