How to Read This Investigation

Investigate decisions. Do not collect features.

This is not a tutorial, reference, or certification guide. We begin with incomplete information, build the simplest reasonable design, let reality expose its weaknesses, and only then name the architecture that survives.

Begin with curiosity

  • Ask what problem forced the design to exist.
  • Prefer a durable question to a memorized answer.

Reason with evidence

  • Separate documented history from architectural derivation.
  • Treat trade-offs and simpler alternatives honestly.
What problem forced engineers to invent this?
Investigation 001 · Core Control Plane Principles

Why a Process
Cannot Keep Itself Alive

"Nothing that can die should be trusted to keep itself alive."

⌄
The Setup

10,000 restaurants. Exactly 5 chefs each.

One morning, chefs call in sick in Mumbai, quit in Delhi, and get rained out in Bengaluru — while extra chefs mistakenly show up in Pune. Some kitchens are understaffed. Some are overstaffed. Who is responsible for noticing, and fixing it — forever, without anyone watching?

👨‍🍳 Chefsthe workers
🏪 Restaurant Managerwatches one kitchen
🏢 Head Officedecides "5 per kitchen"
A stability owner notices the kitchen has too few chefs and restores the requested count. A strategy owner decides when the menu itself should change.
The Question That Starts Everything
"If one supervisor already knows what should run, why not let it own everything directly?"

This investigation answers that question the hard way — by building the simplest possible design, and then letting reality break it, one small crack at a time.

Part I · The Simplest Idea

One component. Does everything.

The obvious design: one supervisor stores the requested outcome, counts the instances that exist, and creates or removes the difference.

Supervisorstores the request
Instancesthe actual work
Read Request
→
Count Instances
→
Compare Desired vs Actual
→
Create / Remove
→
Repeat

The supervisor repeats the same four moves. For now, we deliberately assume it can observe the count.

0
Running Instances
"A successful demo proves almost nothing."
The Demo That Fooled Everyone

It worked. The room clapped.

Then one quiet engineer asked: "What responsibilities does the supervisor actually own?" The team started listing them out loud — and didn't stop.

Creates InstancesMaintains requested countRolling changes RollbacksRelease historyStatus reporting Progress trackingOwnership...future features
"Architecture rarely becomes complicated overnight.
It happens one responsibility at a time."
Part II · Breaking Our Design

Eight small cracks. One pattern.

None of these incidents were catastrophic on their own. Each one just asked a slightly harder question than the last.

01

One Instance Disappeared

The request still says five. Reality now contains four. The original create action has already finished.

Creating work and preserving an outcome are different responsibilities.

02

The Supervisor That Knew Too Much

Replica count, rollout strategy, history, rollback, status, and progress all accumulate in one component.

Nothing is broken yet, but every new reason to change increases architectural coupling.

03

When Responsibilities Collide

Keeping today stable and moving from one version to another demand different decisions at different times.

Responsibilities that change for different reasons need different owners.

04

Who Owns the Instances?

Several supervisors and many instances now share one system. Counting everything can no longer identify responsibility.

Creation is an event. Ownership is the continuing relationship that governs what happens next.

05

The Update That Changed Everything

Overwriting the old version destroys the state needed to reverse a failed change.

A new stability group per version preserves rollback as a reverse rollout.

06

The Machine That Disappeared

Several instances vanish together. The stability responsibility sees a shortage, not a trustworthy explanation.

The owner of replica count must correct drift without absorbing infrastructure diagnosis.

07

The Ownership Crisis

The strategy owner is deleted. The system must know which stability group and instances share its lifecycle.

Lifecycle behavior requires explicit ownership, not historical memory of who created what.

08

The Split Becomes Inevitable

Every honest question leads to the same boundary: strategy above, stability below, executable work at the edge.

The extra component is justified by fewer reasons for confusion, not by a preference for more boxes.

Compact review
The Turning Point

The Responsibility Split

"If two responsibilities change
for different reasons —

they should not live in the same component."

Only now do we give the discovered roles their Kubernetes names.

Deploymentdeployment strategy
ReplicaSetreplica management
Podsthe actual work
↓ keep scrolling — watch the responsibility split ↓

Deployment gets simpler — it no longer counts or recreates Pods. The new controller has exactly one job: maintain the number of replicas requested. It didn't even have a name for the first twenty minutes. The idea mattered before the name did.

A note on history. This investigation reasons outward from the architectural problem, so its derivation sequence differs from Kubernetes' historical evolution. The real path ran from bare Pods to ReplicationController and later to Deployments and ReplicaSets. The principles are the same; this teaching path makes the necessity of each responsibility visible.
Interactive Architecture Lab

Make the boundary do real work.

Disturb count, request a new version, and watch which responsibility must act. The simulation is deliberately bounded: it demonstrates ownership, not how observations arrive.

Requested count: 3Requested version: v1
Instances
3 / 3
Running version
v1
Last owner
Combined supervisor

Reality matches the request.

Creation Is an Event. Ownership Is a Relationship.

Who does this Pod actually belong to?

A carpenter builds a chair. A customer owns it. The carpenter's job ended at creation — the customer's relationship to the chair keeps going: who repairs it, who decides to throw it away.

Deployment
owns
ReplicaSet
owns
Pods

Delete the Deployment, and this chain is exactly what tells the cluster what should happen next — not just who created what, but who's responsible for what happens after. ownerReferences didn't invent this idea; it just records an architectural necessity that already existed.

Why Not Just Update in Place?

Because rollback needs history that still exists.

Overwrite the ReplicaSet's Pod template for v2, and there's nothing left to roll back to. Instead: a new ReplicaSet per version. A rollout is just two numbers moving in opposite directions.

ReplicaSet v1
5
ReplicaSet v2
0
↓ scroll to run the rollout ↓

v1 is scaled to zero and retained as bounded rollout history. It becomes a temporary time capsule until revision-history limits reclaim older ReplicaSets. Rollback isn't instant magic; it's the same rollout, run in reverse.

The Single Biggest Insight
Responsibilities that change
for different reasons
deserve different homes.
Not fewer boxes on the diagram. Fewer reasons for confusion.

Deployment

Architect of change
  • What version should run?
  • How do we move v1 → v2?
  • How many Pods can be down mid-update?
  • Should this rollout pause or roll back?

ReplicaSet

Guardian of stability
  • Does observed count match desired count?
  • Doesn't care about versions
  • Doesn't care about rollout timing
  • Observe. Compare. Correct. That's it.
What responsibility does this component own? +
Ask this before extending any controller — if the answer takes more than one sentence, it probably owns too much.
What has it deliberately refused to own? +
ReplicaSet refuses to know why a Pod disappeared, what node it ran on, or how scheduling works. That refusal is the design.
If it disappeared tomorrow, where would its job go? +
If the answer is "nowhere obvious," that's a sign the responsibility was never clearly owned in the first place.
Architectural Honesty

Keep the direct supervisor

  • One small system and release path
  • No independent scaling policy
  • No retained rollback history required
  • The extra boundary would cost more than it clarifies

Split the responsibilities

  • Stability and rollout change independently
  • Ownership must outlive individual actions
  • Rollback requires retained version history
  • Each component needs one defensible reason to change
Costs Accepted

Component

Another active responsibility must be operated and understood.

Relationship

Ownership and lifecycle now cross an explicit boundary.

State

Version history and desired counts must remain coherent.

Delay

Correction is eventual, not instantaneous; observation remains unresolved.

This Isn't Just Kubernetes

The same split, everywhere.

Flip each card — the same principle, wearing a different name.

Investigation Exercise

Prediction before command.

The point is not the syntax. It is whether you can predict behavior from ownership.

01 · Prediction

Commit before observing

Will a bare Pod return after deletion? Will work with an external owner return?

02 · Experiment

Run one sequence

Create and delete both forms of work, then ask the cluster what remains.

03 · Observation

Compare outcomes

The bare Pod remains absent. The managed Pod is replaced with a new identity.

04 · Reflection

Name the difference

Replacement comes from an external owner preserving an invariant, not from the failed process.

One Question We Never Answered

How does a ReplicaSet actually
know reality changed?

Did it ask the API Server?

Does it poll — every second? Every minute?

What if there are a million Pods?

What if a thousand disappear at once?

Observation boundary. This companion says the stability owner “notices” drift. That is deliberate shorthand. A distributed system sees delayed reports, not a globally current present. How change is discovered and how observations become trustworthy are later mysteries.
Ownership explains who must preserve the invariant. It does not explain how that owner learns reality changed.
Next Investigation

INV-002 — The Reconciliation Pattern

One of the most universal patterns in computing — showing up in databases, operating systems, robotics, and control theory, not just Kubernetes.

Deliberate Simplifications Ledger

How the replica count is actually kept correct over time (the loop itself)INV-002 · The Reconciliation Pattern
Who runs that loop, and why there are many such workersINV-003 · The Controller Pattern
The mechanism that records ownership beyond the template fingerprint described hereINV-012 · Owner References
What actually happens to owned objects when an owner is deleted (the cascade)INV-013 · Garbage Collection
How many past versions are retained and how old ones are reclaimedINV-013 · Garbage Collection

Key Terms

Desired State
What you asked for — "always 5 running" — not a sequence of steps to get there.
Observed / Actual State
What's currently true in the cluster right now.
ReplicaSet
The controller born from splitting "keep N running" away from deployment strategy.
Ownership
An ongoing relationship — distinct from creation — that governs lifecycle and deletion.
Reconciliation
Continuously comparing desired vs. observed state and correcting the difference. (Full story: INV-002.)
Time Capsule
An old ReplicaSet scaled to zero and retained within bounded rollout history, so a retained revision can be scaled back up.
Design Discomfort
The feeling that something's wrong architecturally, even though nothing is actually broken yet.