Resources are finite
CPU time and memory are bounded physical capacities, not elastic promises.
Investigation 019 · Node Execution Principles
We inherit two earned contracts without rediscovery: the Pod is the execution unit, and a Pod has its own network identity. The mystery here is different: if independent Pods share one machine, why can they not all consume resources without explicit ownership?
Begin the investigation ↓Prologue
A node begins with a small set of independent workloads. Nothing appears unfair. Every process advances and the machine remains responsive.
Platform checks only whether a workload can start right now.
After start, workloads use CPU and memory without explicit ownership boundaries.
Light demand makes passive sharing look rational and sufficient.
Collectively rising demand turns silent assumptions into architectural failure.
First Principles
A single workload can consume everything harmlessly. Two independent workloads cannot, because they optimize independently against the same finite hardware.
CPU time and memory are bounded physical capacities, not elastic promises.
Workloads optimize for their own outcomes, not for cluster-wide fairness.
When total demand can exceed supply, allocation decisions become unavoidable.
Without ownership, allocation is accidental. Without allocation, protection is accidental.
Naive Architecture
If yes, the workload starts. After that, each workload consumes CPU and memory as it chooses. No explicit ownership, no declared boundaries, no pre-scarcity contract.
The Architecture That Almost Worked
Early workloads are small and bursty. Demand stays below capacity, so unmanaged sharing looks like a principled design rather than a temporary coincidence.
Low steady CPU, modest memory.
Intermittent batch spikes under normal limits.
Memory grows slowly with demand.
Assumes independent workloads remain collectively rational forever.
Breaking Our Design
Pressure. One workload can abruptly increase CPU demand while unrelated workloads keep normal behavior.
Prediction. If passive sharing is enough, unrelated progress should remain stable as one neighbor becomes greedy.
Adjust one workload's demand on a 4-core node.
Unrelated workload progress degrades as greedy demand rises.
No owner can protect independent progress.
CPU sharing needs explicit ownership boundaries.
Memory scarcity is harsher than delay.
This is an illustrative contention model, not OS scheduling simulation.
Progress cannot depend on neighbor restraint.
This experiment models demand versus finite capacity. It does not claim to simulate real kernel scheduling.
Prediction: unrelated progress remains stable while one neighbor demands more CPU.
Pressure. Individually plausible memory growth can collectively exceed finite physical capacity.
Prediction. If no ownership exists, memory pressure should still resolve predictably.
Use explicit 16 GiB node and three plausible allocations.
Total physical memory demand exceeds 16 GiB.
An unspecified loser is forced without platform policy.
Memory ownership must be explicit before scarcity.
Can cooperation replace ownership?
No deterministic victim or OOM mechanism is claimed here.
Admission alone is insufficient for survival guarantees.
Prediction: plausible allocations remain inside physical memory.
Pressure. Every participant can promise restraint, yet future demand remains unknowable and uncoordinated.
Prediction. Mutual pledge should keep fairness stable without platform ownership.
All workloads pledge conservative usage.
Unpredictable surge or leak breaks assumptions.
Pledges are unverifiable and coordination collapses.
Fairness must be platform-owned, not participant-promised.
Which responsibilities must the platform own?
No participant has global foresight of future demand.
Trust cannot replace explicit pre-scarcity rules.
Prediction: voluntary cooperation remains sufficient under changing demand.
Pressure. Repeated failures show missing ownership of allocation and protection responsibilities.
Prediction. If responsibilities are explicit before scarcity, the platform can make later allocation and protection decisions coherently.
Classify responsibilities as unowned or platform-owned.
Owned model identifies who must define allocation and protection before pressure.
Ownership alone does not enforce protection; mechanism remains required.
Ownership, allocation, and protection are platform contracts.
How should a contract be expressed generically?
This episode models ownership only; no cgroup or throttling enforcement is implemented.
Policy ownership is distinct from enforcement mechanism.
Prediction: unowned and owned models are equivalent under scarcity pressure.
Compact Review
The Turning Point
A passive platform can start workloads. A reliable platform must own allocation and protection responsibilities before contention begins, even though enforcement remains unresolved.
Generic Resource Ownership Contract
Finite resources require explicit ownership, allocation, and protection responsibilities before scarcity appears.
CPU and memory are bounded capacities that must be treated as shared assets.
Define ownership and admissible claims before overload turns every decision into an emergency.
One workload's behavior must not arbitrarily erase another workload's predictability.
When demand exceeds supply, decisions must follow explicit policy boundaries.
Ownership and allocation rules are architectural policy; enforcement is mechanism and may vary.
Capacity, workload shape, and demand shift over time; ownership requires ongoing maintenance.
| Responsibility | Unowned Model | Owned Model |
|---|---|---|
| Capacity recognition | implicit and reactive | explicit and continuous |
| Allocation decisions | emergent under pressure | defined before pressure |
| Workload protection | best effort by accident | deliberate contract obligation |
| Policy versus mechanism | collapsed implicitly | separated intentionally |
Only Now: Kubernetes
Kubernetes resource requests contribute accounted claims used by scheduling. Optional limits provide enforcement inputs. Observed consumption is delayed evidence, and physical capacity is the machine's actual finite resource. These states are related but not interchangeable.
Scheduler accounting, kubelet and kernel enforcement paths, QoS behavior, and eviction policy are mechanisms and policy choices. They operationalize the contract but do not redefine the architectural requirement.
Physical capacity ≠ accounted claim ≠ enforcement input ≠ observed consumption.
Engineering Reflection
The investigation's timeless chain is direct: Ownership → Allocation → Protection → Reliable Execution.
Finite resources need explicit owners before demand conflict.
Ownership enables consistent decisions when demand exceeds supply.
Allocation boundaries shield unrelated workloads from collateral instability.
Protected workloads preserve progress despite neighbor pressure.
When workloads are few, operators are close, and failures are acceptable, informal cooperation can seem workable.
As independence and scale rise, fairness by trust becomes unverifiable and eventually unsafe.
Platform must continuously track finite shared resources.
Allocation rules must be explicit and reviewable.
Independent workloads require boundary-aware decisions.
Ownership must evolve with demand and workload changes.
Investigation Exercise (Optional)
Predict each outcome before running it. Cancel midway to inspect partial evidence.
Increase one workload's demand on a 4-core model.
Expand plausible allocations beyond 16 GiB total.
Apply a fairness pledge, then introduce unpredictable demand.
Classify responsibilities and compare outcomes.
Bridge to INV-020
INV-019 established who must own allocation and protection responsibilities on one node. It did not choose an economic objective or implement enforcement, and it does not explain cross-node communication.
We can protect finite resources on one node. We still must make workloads reachable across nodes.