Investigation 013 · The Garbage Collection Problem
An owner disappears.
What happens to everything it owned?
An Application owns Deployments, which own ReplicaSets, which own Pods. Ownership tells the system who created what. It says nothing about what should happen when the object at the top of that chain is deleted.
Begin the investigation ↓Prologue
Creating resources is easy. Disappearing is the hard part.
A web server writes temporary logs. A database creates snapshots. A cloud platform provisions virtual machines, storage volumes, load balancers. Creation is common, cheap, and safe to get slightly wrong — recreate it and move on.
Deletion is different. An application creates a database, which creates storage volumes, which allocate physical disks. Delete the application — should every dependent resource vanish automatically? Probably. Now two applications share that same database, and one is deleted. Should the database disappear now? Probably not.
Already, deletion depends on more than ownership. Add scale — thousands of resources created per minute, hundreds deleted simultaneously, controllers restarting mid-operation, networks partitioning — and every deletion becomes an architectural decision. Delete too aggressively and healthy resources disappear. Too cautiously, and abandoned resources accumulate forever. In the wrong order, and dependencies break.
First Principles
Creation is a local decision. Deletion is a global one.
“Delete it, it’s no longer needed” sounds obvious — but a distributed system doesn’t understand need, intent, or importance. It only understands state. Every resource exists for a reason: some are created directly, others — derived resources — exist only because something else required them. A Pod exists because a ReplicaSet required it; the ReplicaSet because a Deployment required it.
Creating a resource requires only one decision-maker: a controller decides it needs an object, and creates it. Deletion is different. Removing a resource changes the state of the entire system — is another component still using it? Does something else depend on it? Will deletion leave the system inconsistent? Deletion cannot always be decided in isolation.
Creation adds possibilities; a mistaken object can usually be removed later. Deletion removes them — often permanently. Recovering a wrongly deleted resource may require backups, manual intervention, or may simply be impossible. This asymmetry is why distributed systems treat deletion with far more caution than creation.
This asymmetry is why distributed systems treat deletion with far more caution than creation.
Naive Architecture
The Lifecycle Chain
Many resources are derived: they exist only because something else required them. A natural rule follows immediately — if a resource exists only because another resource exists, it should disappear once that reason disappears.
Delete the Application, and the rule cascades cleanly downward: the Deployment goes, then the ReplicaSet, then the Pods. No abandoned resources, no manual cleanup — the system appears to clean itself automatically.
Lab — Cascade a Deletion Down the Chain
Delete the Application and watch the “delete everything beneath it” rule propagate down the lifecycle chain.
Every object exists because of its parent. Nothing has been deleted yet.
This is our first architectural hypothesis: if a resource exists because of another, deleting the parent should automatically delete every dependent resource. It is elegant, deterministic, and easy to automate. It is also about to be tested.
It is elegant, deterministic, and easy to automate. It is also about to be tested.
The Architecture That Almost Worked
A self-cleaning system — with one comforting assumption
The cascading-delete rule has real virtues. It is deterministic — every deletion has a predictable outcome. It is automatic — nobody has to remember to clean up dependents. It is scalable — the same rule applies to ten resources or ten million. Responsibility stays local: each object only needs to know its immediate children.
It appears to solve one of the most persistent operational problems in large systems: resource leakage. Temporary files, unused virtual machines, orphaned storage volumes, network rules nobody remembers creating — all of it seems to vanish automatically under this one rule.
This is one of those rare moments in engineering where a design feels finished. But detective stories get interesting exactly when the evidence looks complete. Reality has not yet been consulted.
Detective stories get interesting exactly when the evidence looks complete.
Breaking Our Design
Every deletion rule fails a different way
Objects are shared. Order matters. Failures happen mid-operation. Under these conditions a single universal deletion rule stops being safe. Four increasingly hostile experiments attack four different assumptions.
Episode 1 — Delete Everything
An application creates a Deployment and, separately, a Configuration object holding tuned parameters another team now depends on. Delete the application, and the cascade rule removes the Configuration too — correctly, according to the rule, and disastrously in practice.
The ownership relationship tells us who created the Configuration. It never told us whether it should survive. Creation history is not lifecycle policy.
Episode 2 — Delete Nothing
Reacting to the first disaster, flip the rule: deleting a parent removes only the parent, and every dependent stays untouched. Nothing valuable is ever destroyed by accident.
Week after week, orphans accumulate — Deployments, ReplicaSets, Secrets, volumes nobody remembers creating. The system becomes a warehouse of forgotten state. We traded deleting too much for deleting too little.
Episode 3 — The Order Problem
Suppose the system finally knows exactly which resources should disappear. It deletes the Deployment before its ReplicaSet has finished disappearing. For a window of time, a ReplicaSet exists whose parent has already vanished.
Another controller observing that intermediate state has no way to know it is temporary. Intermediate states are real states — correctness must hold throughout deletion, not only after it completes.
Episode 4 — The Partial Failure Problem
Deletion begins in the correct order. Halfway through, a controller crashes. Some resources are gone; others remain. When the controller restarts, does it start over, skip ahead, or guess?
If cleanup progress lived only in that controller’s memory, it is gone. Deletion must be resumable — its progress must live in the system, not inside a single process.
Deletion must be resumable — its progress must live in the system, not inside a single process.
Lab — Test Every Deletion Policy
Try each universal deletion rule against a realistic scenario. Watch each one collapse.
Click a policy to test it against reality.
Turning Point
We have exhausted every universal deletion rule
Deleting everything destroys resources someone still needs. Deleting nothing buries the system in orphans. Ignoring order exposes broken intermediate states to every observer. Ignoring failure loses cleanup progress the moment a controller crashes.
Ownership alone cannot decide any of this. It records relationships — nothing more. If lifecycle cannot be derived from ownership, it must become its own explicit architectural contract.
If lifecycle cannot be derived from ownership, it must become its own explicit architectural contract.
The Garbage Collection Contract
Five responsibilities every correct design must satisfy
Never assume every dependent should be deleted — different resources may legitimately require different outcomes.
Ownership is evidence, not a deletion command. It must be interpreted under current policy before anything is removed.
Every intermediate state visible to other components must remain architecturally valid throughout deletion.
Progress cannot live only in a controller’s memory — lifecycle state must be durable and observable by the system itself.
Temporary failures must never leave resources permanently abandoned. Garbage collection is a convergence process, like reconciliation.
Kubernetes satisfies this contract through Owner References plus a Garbage Collector controller that evaluates the ownership graph and applies lifecycle policy — deciding whether to cascade, preserve, or delay:
ownerReferences:
- apiVersion: apps/v1
kind: ReplicaSet
name: frontend-rs
uid: 8f3c...
controller: true
blockOwnerDeletion: trueLab — Cleanup Must Survive Failures
Advance a resource through its lifecycle states, then crash the controller mid-cleanup. Does progress survive?
Lifecycle state: Active. Ready to begin deletion.
Engineering Reflection
Relationships describe the world. Policies determine how it changes.
Creating resources is easy. Removing them safely is one of the hardest problems in distributed systems — deletion is not creation in reverse, it is a fundamentally different engineering problem.
Ownership gives structure. Garbage collection governs behavior. Confusing the two produces systems that are either dangerously aggressive or permanently conservative. Correctness must also hold throughout every intermediate state a deletion passes through, not merely at its conclusion — every controller and every API client can observe those states and act on them.
Costs Accepted
Additional controllers. A dedicated garbage collector evaluates the ownership graph and enforces policy — a responsibility ownership alone never had.
Asynchronous deletion. Cleanup may take multiple reconciliation cycles instead of one atomic operation.
Explicit lifecycle states. Active, scheduled, waiting on dependencies, partially cleaned up, ready for removal — every stage must be visible, not just the start and end.
Durable progress tracking. Crash recovery requires knowing exactly what already happened, not just what should happen next.
These costs are intentional — significantly smaller than the operational cost of accidentally deleted state, permanently orphaned resources, or operators reconstructing lifecycle decisions by hand.
Investigation Exercise
Exercise 1 — Delete everything. Build a small dependency graph and delete the root, cascading automatically down every link. Which removed resources might another component still have needed?
Exercise 2 — Delete nothing. Reset the graph, delete only the root, leave every dependent untouched. What orphans accumulate, and how would anyone know they are safe to remove?
Exercise 3 — Change the order. Delete a parent before its children finish disappearing. Could another controller observe the broken intermediate state and make an incorrect decision?
Exercise 4 — Simulate a crash. Stop cleanup halfway through, then resume. What information had to survive the crash for the system to continue correctly?
Bridge to INV·014
Knowing what should disappear is not the same as knowing what should run
Ownership and lifecycle are now both explicit. The control plane knows desired state, who owns what, how to reconcile differences, and how resources should safely disappear.
But suppose a Deployment reconciles down to a Pod object that should exist. Nothing runs yet — no container starts, no process executes. The Pod exists only as a record in the system’s shared source of truth. The system knows what should exist. It has not decided where.
The next investigation explores that placement decision — the boundary between deciding what the system should look like and deciding how that decision reaches a physical machine.
Deliberate Simplifications Ledger
| We deliberately postponed | Owned by |
|---|---|
| Where a Pod should actually execute once it exists as an object | INV·014 — The Scheduling Problem |
| Deletion timestamps and the finalizer mechanism in detail | Implementation detail — metadata.deletionTimestamp / metadata.finalizers |
| Foreground vs. background vs. orphan deletion propagation policies | Implementation detail — DeleteOptions.propagationPolicy |
| Resources shared across multiple owner references | Implementation detail — ownerReferences slice / reference counting |
| The garbage collector’s graph-building and dirty-queue reconciliation loop | Implementation detail — pkg/controller/garbagecollector |
Sources
Official Documentation: Kubernetes Documentation — Garbage Collection; Kubernetes Documentation — Owners and Dependents; Kubernetes Documentation — Finalizers; Kubernetes API Conventions — deletion and lifecycle fields.
Source Code: pkg/controller/garbagecollector — the garbage collector implementation (graph builder + dirty queue); k8s.io/apimachinery/pkg/apis/meta/v1 — DeleteOptions.PropagationPolicy.
Next: INV·014 — The Scheduling Problem