Execution identity
One Pod instance address for one temporary run.
Investigation 021 - Cluster Networking Principles
We inherit one contract without rediscovery: from INV-020, every Pod can directly reach every other Pod on a flat cluster network. We also inherit one limit: a Pod IP belongs to one temporary Pod instance. If that instance disappears after lookup, the old address may still be reachable and still be empty.
Begin the investigation downPrologue
Orders calls payment at Pod 10.42.7.19 and succeeds all day. The payment Pod crashes, replacement starts at 10.42.13.8, and orders still calls 10.42.7.19. The network keeps its promise and routes correctly. Nothing listens.
Client learns the current Pod address.
Pod replacement happens between lookup and connect.
Client uses now-stale address.
Failure appears without violating network reachability.
First Principles
If communication binds to the shortest-lived identity, callers inherit every replacement event. If communication binds to longer-lived capability intent, replacement can be absorbed elsewhere.
One Pod instance address for one temporary run.
The capability callers expect to remain while implementations change.
The caller wants payment processing, not whichever Pod happened to exist at one instant.
Distributed systems continue changing between observation and action. That timing gap is irreducible.
Naive Architecture
Clients avoid long-lived caches by asking the registry every time. The registry reports current Pod address truth at lookup time.
The Architecture That Almost Worked
After Pod replacement, a subsequent lookup can return the new Pod address and recover quickly. But lookup and connect are separate events in time, and the gap between them cannot be removed.
Clients no longer pin one Pod IP forever.
With infrequent replacement and cheap retries, this can be acceptable.
Registry truth at lookup time is not a guarantee at connect time.
Breaking Our Design
Pressure. Client receives a current Pod address, then Pod dies before connect.
Prediction. If lookup freshness is enough, connection should still succeed.
Run lookup, then terminate Pod before connect.
Timeline shows truthful lookup and failed connect.
Address truth does not imply lifetime continuity.
Lookup truth is point-in-time truth only.
What changes when one service has many replicas?
No health policy is derived here.
Timing gap is irreducible.
Prediction: lookup freshness should preserve successful connection.
Pressure. Service adds replicas and registry returns multiple Pod addresses.
Prediction. A registry returning several current addresses should be enough for callers to communicate without assuming another responsibility.
Return several replica addresses and require the caller to choose one.
Every selection strategy adds traffic-distribution responsibility to the caller.
The registry reports membership but does not own request destination choice.
Registry discovery and traffic distribution are different responsibilities.
What happens when many clients do this independently?
No policy winner is declared.
Replica choice is not solved by lookup alone.
Prediction: several current addresses are sufficient without adding client coordination responsibility.
Pressure. Many clients each track membership, infer eligibility, choose replicas, and retry.
Prediction. Independent client views should remain synchronized enough to behave as one view.
Compare one platform view against multiple client-local snapshots after churn.
Clients refresh at different times and hold delayed, partial views.
Coordination duplicates across clients and diverges.
Platform knowledge is reconstructed many times from lagging signals.
Can one stable identity absorb this once?
Eligibility and readiness policy are still unresolved.
Delayed observation must be acknowledged explicitly.
Prediction: client and platform views should remain effectively identical.
Pressure. Callers need stable intent while Pod topology changes continuously.
Prediction. If callers own topology details, one more client improvement should be enough.
Assemble contract guarantees at architecture level only.
Each missing guarantee reproduces earlier pressures.
Complete set closes all four episodes.
Stable service identity and platform mapping are required.
How do names resolve to that identity? INV-022.
No proxy or data-plane implementation is modeled here.
Contract is earned before mechanism.
Prediction: one more client-side improvement should close the problem.
Optional Episode Review
The Turning Point
The architecture shifts from per-client topology reconstruction to one platform-owned, continuously reconciled mapping between stable service identity and current execution instances.
Generic Stable Address Contract
A stable service identity must outlive Pod lifetimes. The platform must own and continuously reconcile mapping to current instances while callers express intent rather than topology.
Service communication identity is independent of any individual Pod lifetime.
Mapping from stable identity to current service instances is platform-owned state.
Mapping follows replacement and scaling continuously, with non-zero propagation windows.
Applications request service capability, not direct replica addresses.
Platform solves coordination once rather than forcing every caller to reconstruct it.
This contract does not claim zero failures, perfect simultaneity, or globally current truth at every instant.
Only Now: Kubernetes
An ordinary virtual-address Service provides a stable address for a capability while Pods behind it can change. Headless and ExternalName variants intentionally realize different contracts and remain outside this investigation.
The mapping to current Pods can be represented by EndpointSlice data as one realization of the generic contract.
This chapter intentionally does not explain how packet forwarding or traffic algorithms are implemented.
Engineering Reflection
When capabilities outlive implementations, callers should remain stable while the platform absorbs topology churn.
One replica, infrequent replacement, and cheap retries can justify direct registry lookup.
As churn and caller count grow, duplicated client coordination exceeds simple lookup benefits.
Platform needs explicit service-to-instance membership declaration.
Eligibility and topology signals are observed over time, not instantaneously.
Mapping updates are continuous but not atomic across all participants.
A platform data-plane component is required even though this chapter does not implement it.
That is 500 partial membership facts reconstructed at the edge of applications, versus one platform-owned service membership view.
No. It only derives that mapping needs eligibility input. Policy details remain open.
Investigation Exercise (Optional)
Predict which ownership model changes when Pod replacement occurs.
Run four-step trace from lookup timing to platform-owned stable identity.
Record where delayed or partial observation changes client behavior.
State what was proven as architecture and what remains mechanism or policy.
Bridge to INV-022
Service identity can now outlive Pod replacement.
Clients no longer need to track every replica address directly.
Platform owns the changing mapping behind stable service identity.
But applications are configured with names, not raw addresses.
The next mystery asks who turns names into current service identity.
Given a service name, who knows the current address now?