Investigation 019 · Node Execution Principles

A Pod has identity and execution.
Who owns finite CPU and memory?

We inherit two earned contracts without rediscovery: the Pod is the execution unit, and a Pod has its own network identity. The mystery here is different: if independent Pods share one machine, why can they not all consume resources without explicit ownership?

Begin the investigation ↓
Node CapacityFINITE · SHARED · CONTESTED

Author's Note

Keep inherited boundaries intact.

This investigation does not reopen Pod identity or machine-local execution ownership. Those contracts are inherited from INV-018 and INV-015 through INV-017.

We only discover resource ownership: how a shared machine remains predictable when independent workloads demand finite CPU and memory.

Execution asks where and whether work runs. Resource ownership asks how much of a finite machine each workload may claim.

Prologue

Calm load hides a missing owner.

A node begins with a small set of independent workloads. Nothing appears unfair. Every process advances and the machine remains responsive.

Admission

Platform checks only whether a workload can start right now.

Execution

After start, workloads use CPU and memory without explicit ownership boundaries.

Observation

Light demand makes passive sharing look rational and sufficient.

Pressure

Collectively rising demand turns silent assumptions into architectural failure.

First Principles

Shared finite capacity creates unavoidable decisions.

A single workload can consume everything harmlessly. Two independent workloads cannot, because they optimize independently against the same finite hardware.

Property 1

Resources are finite

CPU time and memory are bounded physical capacities, not elastic promises.

Property 2

Goals are independent

Workloads optimize for their own outcomes, not for cluster-wide fairness.

Property 3

Competition emerges

When total demand can exceed supply, allocation decisions become unavoidable.

Without ownership, allocation is accidental. Without allocation, protection is accidental.

Naive Architecture

Admission asks only: is there room right now?

If yes, the workload starts. After that, each workload consumes CPU and memory as it chooses. No explicit ownership, no declared boundaries, no pre-scarcity contract.

Admission Checkstart if room exists now
Unrestricted Runtimeuse CPU and memory freely
Passive Platformobserve after the fact

The Architecture That Almost Worked

Under light load, passive sharing appears sound.

Early workloads are small and bursty. Demand stays below capacity, so unmanaged sharing looks like a principled design rather than a temporary coincidence.

Workload A

Web API

Low steady CPU, modest memory.

Workload B

Background Worker

Intermittent batch spikes under normal limits.

Workload C

Cache

Memory grows slowly with demand.

Hidden premise

Collective restraint

Assumes independent workloads remain collectively rational forever.

Breaking Our Design

Four causal episodes derive resource ownership.

EPISODE 01

CPU Greedy Neighbor

Pressure. One workload can abruptly increase CPU demand while unrelated workloads keep normal behavior.

Prediction. If passive sharing is enough, unrelated progress should remain stable as one neighbor becomes greedy.

Experiment

Adjust one workload's demand on a 4-core node.

Observation

Unrelated workload progress degrades as greedy demand rises.

Failure

No owner can protect independent progress.

Discovery

CPU sharing needs explicit ownership boundaries.

Next Pressure

Memory scarcity is harsher than delay.

Constraint

This is an illustrative contention model, not OS scheduling simulation.

Boundary Delta

Progress cannot depend on neighbor restraint.

Local model: 4-core node with one adjustable greedy workload.

This experiment models demand versus finite capacity. It does not claim to simulate real kernel scheduling.

Greedy WorkloadDemand: 1.0 cores
Unrelated WorkloadProgress: stable
Total Demand2.0 / 4.0 cores

Prediction: unrelated progress remains stable while one neighbor demands more CPU.

EPISODE 02

Memory Exhaustion

Pressure. Individually plausible memory growth can collectively exceed finite physical capacity.

Prediction. If no ownership exists, memory pressure should still resolve predictably.

Experiment

Use explicit 16 GiB node and three plausible allocations.

Observation

Total physical memory demand exceeds 16 GiB.

Failure

An unspecified loser is forced without platform policy.

Discovery

Memory ownership must be explicit before scarcity.

Next Pressure

Can cooperation replace ownership?

Constraint

No deterministic victim or OOM mechanism is claimed here.

Boundary Delta

Admission alone is insufficient for survival guarantees.

Node memory capacity: 16 GiB. Plausible allocations can still overrun it.

Workload A6 GiB
Workload B5 GiB
Workload C4 GiB

Prediction: plausible allocations remain inside physical memory.

EPISODE 03

Fairness Cannot Be Voluntary

Pressure. Every participant can promise restraint, yet future demand remains unknowable and uncoordinated.

Prediction. Mutual pledge should keep fairness stable without platform ownership.

Experiment

All workloads pledge conservative usage.

Observation

Unpredictable surge or leak breaks assumptions.

Failure

Pledges are unverifiable and coordination collapses.

Discovery

Fairness must be platform-owned, not participant-promised.

Next Pressure

Which responsibilities must the platform own?

Constraint

No participant has global foresight of future demand.

Boundary Delta

Trust cannot replace explicit pre-scarcity rules.

Local model: cooperative pledge, then unpredictable demand change.

All WorkloadsPledge: restrained usage
Unexpected Eventnone yet
Coordination Outcomestable for now

Prediction: voluntary cooperation remains sufficient under changing demand.

EPISODE 04

The Platform Must Intervene

Pressure. Repeated failures show missing ownership of allocation and protection responsibilities.

Prediction. If responsibilities are explicit before scarcity, the platform can make later allocation and protection decisions coherently.

Experiment

Classify responsibilities as unowned or platform-owned.

Observation

Owned model identifies who must define allocation and protection before pressure.

Failure Exposed

Ownership alone does not enforce protection; mechanism remains required.

Discovery

Ownership, allocation, and protection are platform contracts.

Next Pressure

How should a contract be expressed generically?

Constraint

This episode models ownership only; no cgroup or throttling enforcement is implemented.

Boundary Delta

Policy ownership is distinct from enforcement mechanism.

Classify responsibilities before scarcity.

Unowned Modeladmit then hope
ResponsibilityChoose owner
Owned Modelnot yet earned

Prediction: unowned and owned models are equivalent under scarcity pressure.

Compact Review

The Turning Point

Ownership first.
Scarcity second.

A passive platform can start workloads. A reliable platform must own allocation and protection responsibilities before contention begins, even though enforcement remains unresolved.

Generic Resource Ownership Contract

Architecture before mechanism.

Finite resources require explicit ownership, allocation, and protection responsibilities before scarcity appears.

Contract 1

Recognize finite resources

CPU and memory are bounded capacities that must be treated as shared assets.

Contract 2

Establish ownership pre-scarcity

Define ownership and admissible claims before overload turns every decision into an emergency.

Contract 3

Protect independent workloads

One workload's behavior must not arbitrarily erase another workload's predictability.

Contract 4

Make allocation explicit

When demand exceeds supply, decisions must follow explicit policy boundaries.

Contract 5

Separate policy and enforcement

Ownership and allocation rules are architectural policy; enforcement is mechanism and may vary.

Contract 6

Manage continuously

Capacity, workload shape, and demand shift over time; ownership requires ongoing maintenance.

ResponsibilityUnowned ModelOwned Model
Capacity recognitionimplicit and reactiveexplicit and continuous
Allocation decisionsemergent under pressuredefined before pressure
Workload protectionbest effort by accidentdeliberate contract obligation
Policy versus mechanismcollapsed implicitlyseparated intentionally

Only Now: Kubernetes

Kubernetes is one expression of the contract.

Kubernetes resource requests contribute accounted claims used by scheduling. Optional limits provide enforcement inputs. Observed consumption is delayed evidence, and physical capacity is the machine's actual finite resource. These states are related but not interchangeable.

Scheduler accounting, kubelet and kernel enforcement paths, QoS behavior, and eviction policy are mechanisms and policy choices. They operationalize the contract but do not redefine the architectural requirement.

Physical capacity ≠ accounted claim ≠ enforcement input ≠ observed consumption.

Engineering Reflection

Ownership leads to reliable execution.

The investigation's timeless chain is direct: Ownership → Allocation → Protection → Reliable Execution.

Ownership

Finite resources need explicit owners before demand conflict.

Allocation

Ownership enables consistent decisions when demand exceeds supply.

Protection

Allocation boundaries shield unrelated workloads from collateral instability.

Reliable Execution

Protected workloads preserve progress despite neighbor pressure.

Architectural Honesty

Small trusted systems

May tolerate informal sharing

When workloads are few, operators are close, and failures are acceptable, informal cooperation can seem workable.

Growing independent systems

Need explicit ownership

As independence and scale rise, fairness by trust becomes unverifiable and eventually unsafe.

Costs Accepted

Capacity accounting

Platform must continuously track finite shared resources.

Policy clarity

Allocation rules must be explicit and reviewable.

Protection logic

Independent workloads require boundary-aware decisions.

Operational discipline

Ownership must evolve with demand and workload changes.

Investigation Exercise (Optional)

Run the synthesis trace.

Predict each outcome before running it. Cancel midway to inspect partial evidence.

01 · CPU

Increase one workload's demand on a 4-core model.

02 · Memory

Expand plausible allocations beyond 16 GiB total.

03 · Cooperation

Apply a fairness pledge, then introduce unpredictable demand.

04 · Ownership

Classify responsibilities and compare outcomes.

Reflection. After the trace, identify which state was physical capacity, which was observed demand, which responsibilities became platform-owned, and which protection mechanism remains unresolved.
● ● ●
Run the trace after predicting each step.

Bridge to INV-020

One node now has a resource-ownership contract.
A second node creates a communication mystery.

INV-019 established who must own allocation and protection responsibilities on one node. It did not choose an economic objective or implement enforcement, and it does not explain cross-node communication.

From node resource ownership to cluster network communication mystery
We can protect finite resources on one node. We still must make workloads reachable across nodes.

Next Investigation

INV-020 — The Cluster Network Problem

How do independently running workloads communicate when they may execute on different nodes?

Deliberate Simplifications Ledger

How ownership claims are encoded in Kubernetes manifestsOpen backlog
How scheduler accounting evaluates aggregate fitTarget: open backlog
How kubelet and kernel enforce boundaries under CPU pressureTarget: open backlog
How memory pressure resolves through QoS and eviction policyTarget: open backlog
How cross-node Pod connectivity is achievedTarget: INV-020