Investigation 015 · Node Execution Principles

A workload has been placed.
Why is nothing running?

The scheduler selected Machine 42 and recorded that decision in authoritative desired state. No process exists. No CPU is used. The machine remains idle.

Begin the investigation ↓
Machine 42PLACED · IDLE

Author's Note

The scheduling contract is already complete.

INV-014 established who decides where work belongs. That decision is authoritative, versioned, and finished once placement is recorded.

This investigation does not reopen placement. It follows the decision across a different boundary: from durable intent into physical action on one machine.

A correct decision can exist forever without performing a single unit of work.

Prologue

The scheduler finishes. Everyone waits.

The scheduler examined the available machines, evaluated CPU and memory, applied placement policy, and selected Machine 42.

Authoritative desired state now says: Run Workload A on Machine 42.

No process exists.

No memory has been allocated. No CPU cycles are consumed.

Machine 42 remains idle. The placement is still perfectly correct.

The system knows where the workload belongs. Nothing in the architecture owns the transition from that knowledge to running reality.

The mystery

Once a global decision has been made, who is responsible for turning it into local reality?

First Principles

A blueprint cannot construct its building.

Scheduling and execution transform different state, require different evidence, and end at different boundaries.

Global decision

Where should it run?

Requires cluster-wide capacity, constraints, competing workloads, and policy. Its output is placed desired state.

Local action

How does it exist here?

Requires running processes, filesystem state, devices, operating-system events, and direct access to local resources.

Information

Placement changes records

A durable assignment can name a destination. It cannot allocate memory or create a process.

Capability

A machine cannot infer intent

Machine 42 can act locally, but capability alone does not tell it which global intentions belong to it.

Desired State → Placed Desired State is not the same transformation as Placed Desired State → Running Reality.

Naive Architecture

Let the scheduler execute after it places.

The scheduler already chose Machine 42. The shortest design gives it one more step: connect to the machine and start the workload directly.

Schedulerdecides and executes
Machine 42passive remote target
Small fleetA handful of machines.
Reliable linksRemote calls complete clearly.
Few eventsLifecycle work stays rare.
Stable machinesLocal failure is exceptional.

Under those assumptions the design is sensible. One component decides. The same component acts. No new service or machine agent is required.

The Architecture That Almost Worked

Separate one central execution service.

The scheduler returns to one responsibility: placement. A new central service owns execution across every machine.

Schedulerrecords placement
Central Execution Serviceall starts · stops · recovery
Every Machinepassive endpoints
Why it nearly works

Clean component names

The scheduler decides. The execution service executes. Each appears to have one job.

Hidden assumption

Execution is remotely knowable

The service assumes network conversations can replace direct local evidence at every machine.

Breaking Our Design

Four experiments relocate responsibility.

Each episode breaks one assumption, earns one conclusion, and leaves the next question unresolved.

EPISODE 01

The Placement With No Executor

Pressure. Workload A is authoritatively assigned to Machine 42, but the architecture contains no component responsible for making it exist.

Prediction. If placement is itself execution, recording the assignment should create a running process.

Experiment

Record the placement.

Observation

Desired state changes. Machine state does not.

Failure

The intended process remains absent indefinitely.

Discovery

Execution is a separate responsibility.

Still unknown

Who should own it?

Does a placement record perform work?

Desired stateWorkload A unplaced
Machine 42idle
Running realityno process

Prediction: will changing desired state create a process?

Next pressure: a separate execution responsibility is necessary. Can one central service own it for every machine?

EPISODE 02

The Centralized Execution Attempt

Pressure. One service now reaches across the network for every start, stop, inspection, and recovery action.

Prediction. If remote execution is ordinary coordination, one service should remain manageable as the fleet grows.

Experiment

Increase machines and lifecycle traffic.

Observation

Remote work concentrates in one queue.

Failure

Latency and uncertain outcomes accumulate centrally.

Discovery

Execution cannot remain one remote responsibility.

Still unknown

Why does the machine remain passive?

Grow the fleet without growing the central owner.

This illustrative model assumes four remote lifecycle events per machine per minute and a fixed central capacity of 4,000 events per minute. These values expose concentration; they are not Kubernetes limits.

Machines10
Illustrative remote events/min40
Illustrative central capacity/min4,000
Illustrative queue growth/min0

At ten machines, this illustrative model appears comfortable.

Next pressure: execution should move closer to each machine. But how does a machine know which global intentions belong to it?

EPISODE 03

The Silent Machine

Pressure. Machine 42 owns the processors, memory, and local operating-system evidence, yet it does not know that Workload A was assigned to it.

Prediction. If local capability is enough, the machine should infer its work without observing authoritative assignments.

Experiment

Compare global and machine-local views.

Observation

Both views are truthful and disconnected.

Failure

The capable machine remains idle.

Discovery

The machine must discover assignments addressed to itself.

Still unknown

Does discovery once suffice?

What does each side know?

Authoritative viewWorkload A → Machine 42
Machine 42 viewno assignment observed
Resultcapable but uninformed

Two views can both be truthful without being connected.

Next pressure: the machine can now start assigned work. What happens seven minutes later when that work disappears?

EPISODE 04

The Execution Feedback Problem

Pressure. Workload A starts successfully at 09:00 and crashes at 09:07. Desired state still requires it.

Prediction. If execution is a one-time action, successful startup should complete the responsibility.

Experiment

Start once, then terminate the work.

Observation

Desired and actual state diverge.

Failure

No owner notices or repairs the mismatch.

Discovery

Execution is continuous local reconciliation.

Boundary earned

Global decision, local continuous action.

Crash work after a successful start.

Desired stateWorkload A must run
Local realitystopped
Execution modeone-time action

Desired state says run. Local reality is stopped. No action has happened.

Compact Review

The Turning Point

Decide globally.
Execute locally.

The boundary follows knowledge and physical responsibility. The global system places work. Each machine continuously makes only its own assignments real.

The Node Execution Contract

One machine. One bounded promise.

A machine-local execution agent transforms desired state assigned to its machine into running reality.

01 · Observe

Assigned intent

Continuously discover workloads addressed to this machine.

02 · Compare

Desired and actual

Compare assigned state with directly observed local reality.

03 · Create

Missing execution

Make assigned work exist without deciding where it belongs.

04 · Remove

Obsolete execution

Stop local work no longer present in assigned desired state.

05 · Restore

After failure

Repair drift whenever running reality no longer matches intent.

06 · Observe

Local health

Use evidence available at the machine while acknowledging observation delay.

07 · Report

Local facts

Publish observations without making global decisions or claiming present certainty.

Boundary

Refuse placement

Do not rank machines, redefine intent, own authoritative truth, or perform garbage collection.

OwnerInputTransformationOutput
SchedulerDesired stateEvaluate machines and placement policyPlaced desired state
Node execution agentAssignments for one machineObserve, compare, correct, reportRunning reality and local observations

Only Now: Kubernetes

Kubernetes calls the local execution agent the Kubelet.

A Kubelet runs on every node. It observes work assigned to that node, compares desired state with local execution, restores missing work, removes obsolete work, and reports local observations.

The name does not change the contract. Kubernetes is one realization of the boundary we already earned.

The scheduler owns where. The Kubelet owns making that decision real here.

Engineering Reflection

A system that decides globally must execute locally.

Global decisions need a global view. Local execution needs local evidence. Combining them forces one component to live in two incompatible information worlds.

Intellectual Lineage

Borg2015 paper

The Borgmaster scheduled; a borglet on every machine started, stopped, restarted, and reported tasks.

Mesos2011

Central resource allocation remained separate from execution near each worker.

Omega2013

Scheduler coordination evolved while the per-machine execution boundary remained.

Kubernetes2014 onward

The Kubelet inherits the same local responsibility pattern.

Architectural Honesty

Keep it combined

Small, fixed systems

One process may reasonably decide and execute when the fleet is small, links are reliable, local state changes rarely, and one operator can understand every target.

Separate responsibility

Distributed execution

Local ownership becomes necessary when machines fail independently, evidence changes continuously, and execution volume grows with the fleet.

Costs Accepted

Long-lived agents

Every machine runs another system component.

Independent loops

Thousands of local reconcilers must behave consistently.

Delayed reports

The control plane sees reported evidence, not direct present reality.

Asynchronous state

Assignments and execution may temporarily diverge.

Operational surface

Local agents must be deployed, upgraded, and diagnosed.

Accepted trade

Adding a machine adds both capacity and the agent responsible for that capacity.

Investigation Exercise

Build the smallest distributed executor.

Use one scheduler and two inert workers. Predict each state change before running the trace.

01 · Place

Record A → Node 1 and B → Node 2. Confirm that neither workload starts.

02 · Observe

Give each worker access only to assignments bearing its identity.

03 · Reconcile

Let each worker compare its assignments with local reality.

04 · Fail

Terminate A and observe which participant can detect and restore it.

● ● ●
Run the trace after predicting which state changes at each step.

A New Mystery

We know who must execute.
What actually creates the running workload?

Throughout this investigation, creation remained one abstract act. We did not derive process isolation, filesystem preparation, networking, or operating-system-specific execution.

A machine-local execution owner makes placed intent real while process-construction mechanics remain unresolved
Responsibility for execution does not require owning every execution mechanism.

Next Investigation

INV-016 — The Container Runtime Problem

Should the node execution agent implement process creation itself, or delegate it to something specialized?

Deliberate Simplifications Ledger

Execution treated as one abstract act; process-creation mechanics remain unknownINV-016
Isolation and its machine-level consequencesINV-018 / INV-019
Status reporting treated as reliable; silent-node behavior remains unresolvedOpen backlog / INV-011
One execution agent assumed per machineINV-015 extension / INV-016

Sources