Investigation 021 - Cluster Networking Principles

The network can reach Pod addresses.
Can applications survive Pod replacement?

We inherit one contract without rediscovery: from INV-020, every Pod can directly reach every other Pod on a flat cluster network. We also inherit one limit: a Pod IP belongs to one temporary Pod instance. If that instance disappears after lookup, the old address may still be reachable and still be empty.

Begin the investigation down
ReachabilityCAN DELIVER TO THE ADDRESSLIFETIME STILL CHANGES

Author's Note

Preserve inherited reachability, discover lifetime ownership.

This chapter does not reopen flat cluster Pod reachability. That was earned in INV-020.

This chapter asks a narrower question: when a service capability outlives any one Pod instance, who owns the changing mapping between service intent and current Pod instances?

A truthful lookup result is not a lifetime guarantee.

Prologue

A packet arrives at the right address. Nobody is there.

Orders calls payment at Pod 10.42.7.19 and succeeds all day. The payment Pod crashes, replacement starts at 10.42.13.8, and orders still calls 10.42.7.19. The network keeps its promise and routes correctly. Nothing listens.

Lookup

Client learns the current Pod address.

Platform change

Pod replacement happens between lookup and connect.

Connect

Client uses now-stale address.

Result

Failure appears without violating network reachability.

First Principles

Execution identity, service identity, and caller intent have different lifetimes.

If communication binds to the shortest-lived identity, callers inherit every replacement event. If communication binds to longer-lived capability intent, replacement can be absorbed elsewhere.

Identity 1

Execution identity

One Pod instance address for one temporary run.

Identity 2

Service capability identity

The capability callers expect to remain while implementations change.

Identity 3

Caller intent

The caller wants payment processing, not whichever Pod happened to exist at one instant.

Distributed systems continue changing between observation and action. That timing gap is irreducible.

Naive Architecture

Before each connection, query a registry for current Pod IP.

Clients avoid long-lived caches by asking the registry every time. The registry reports current Pod address truth at lookup time.

Clientneed payment now
Registrycurrent Pod IP
10.42.7.19current Pod instance

The Architecture That Almost Worked

A fresh lookup often reaches replacement on the next attempt.

After Pod replacement, a subsequent lookup can return the new Pod address and recover quickly. But lookup and connect are separate events in time, and the gap between them cannot be removed.

Strength

No stale static config

Clients no longer pin one Pod IP forever.

Strength

Simple single-replica operation

With infrequent replacement and cheap retries, this can be acceptable.

Boundary

Timing gap remains

Registry truth at lookup time is not a guarantee at connect time.

Breaking Our Design

Four visible causal episodes derive stable service identity ownership.

EPISODE 01

When the Pod Dies

Pressure. Client receives a current Pod address, then Pod dies before connect.

Prediction. If lookup freshness is enough, connection should still succeed.

Experiment

Run lookup, then terminate Pod before connect.

Observation

Timeline shows truthful lookup and failed connect.

Failure

Address truth does not imply lifetime continuity.

Discovery

Lookup truth is point-in-time truth only.

Next Pressure

What changes when one service has many replicas?

Constraint

No health policy is derived here.

Boundary Delta

Timing gap is irreducible.

Independent simulation: lookup, death, connect.

Lookupnot performed
Pod staterunning
Connectnot attempted

Prediction: lookup freshness should preserve successful connection.

EPISODE 02

Replicas Have No Unified Address

Pressure. Service adds replicas and registry returns multiple Pod addresses.

Prediction. A registry returning several current addresses should be enough for callers to communicate without assuming another responsibility.

Experiment

Return several replica addresses and require the caller to choose one.

Observation

Every selection strategy adds traffic-distribution responsibility to the caller.

Failure

The registry reports membership but does not own request destination choice.

Discovery

Registry discovery and traffic distribution are different responsibilities.

Next Pressure

What happens when many clients do this independently?

Constraint

No policy winner is declared.

Boundary Delta

Replica choice is not solved by lookup alone.

Independent simulation: same replica set, different strategies.

Replica set[10.42.7.19, 10.42.13.8, 10.42.22.5]
Chosen Podnone
Resultnot evaluated

Prediction: several current addresses are sufficient without adding client coordination responsibility.

EPISODE 03

Every Client Becomes a Service Proxy

Pressure. Many clients each track membership, infer eligibility, choose replicas, and retry.

Prediction. Independent client views should remain synchronized enough to behave as one view.

Experiment

Compare one platform view against multiple client-local snapshots after churn.

Observation

Clients refresh at different times and hold delayed, partial views.

Failure

Coordination duplicates across clients and diverges.

Discovery

Platform knowledge is reconstructed many times from lagging signals.

Next Pressure

Can one stable identity absorb this once?

Constraint

Eligibility and readiness policy are still unresolved.

Boundary Delta

Delayed observation must be acknowledged explicitly.

Independent simulation: platform updates once, clients observe later.

Platform-observed membership[A, B, C, D, E]
Client viewsall synchronized
Coordination cost100 x 5 = 500 membership facts

Prediction: client and platform views should remain effectively identical.

EPISODE 04

Derive Stable Service Identity Ownership

Pressure. Callers need stable intent while Pod topology changes continuously.

Prediction. If callers own topology details, one more client improvement should be enough.

Experiment

Assemble contract guarantees at architecture level only.

Observation

Each missing guarantee reproduces earlier pressures.

Success

Complete set closes all four episodes.

Discovery

Stable service identity and platform mapping are required.

Next Pressure

How do names resolve to that identity? INV-022.

Constraint

No proxy or data-plane implementation is modeled here.

Boundary Delta

Contract is earned before mechanism.

Independent simulation: assemble guarantees without implementing packet forwarding.

Stable service identitymissing
Platform-owned mappingmissing
Continuous reconciliationmissing

Prediction: one more client-side improvement should close the problem.

Optional Episode Review

The Turning Point

Callers express service intent.
The platform owns changing execution topology.

The architecture shifts from per-client topology reconstruction to one platform-owned, continuously reconciled mapping between stable service identity and current execution instances.

Generic Stable Address Contract

Guarantees before Kubernetes mechanism.

A stable service identity must outlive Pod lifetimes. The platform must own and continuously reconcile mapping to current instances while callers express intent rather than topology.

Contract 1

Stable service identity

Service communication identity is independent of any individual Pod lifetime.

Contract 2

Platform owns mapping

Mapping from stable identity to current service instances is platform-owned state.

Contract 3

Continuous updates with windows

Mapping follows replacement and scaling continuously, with non-zero propagation windows.

Contract 4

Intent over topology

Applications request service capability, not direct replica addresses.

Contract 5

One coordination solution

Platform solves coordination once rather than forcing every caller to reconstruct it.

This contract does not claim zero failures, perfect simultaneity, or globally current truth at every instant.

Only Now: Kubernetes

Kubernetes commonly realizes the stable identity with a Service.

An ordinary virtual-address Service provides a stable address for a capability while Pods behind it can change. Headless and ExternalName variants intentionally realize different contracts and remain outside this investigation.

The mapping to current Pods can be represented by EndpointSlice data as one realization of the generic contract.

This chapter intentionally does not explain how packet forwarding or traffic algorithms are implemented.

Engineering Reflection

Timeless principle: bind communication to the longer-lived identity.

When capabilities outlive implementations, callers should remain stable while the platform absorbs topology churn.

Architectural Honesty

Where lookup can be enough

Simple low-churn cases

One replica, infrequent replacement, and cheap retries can justify direct registry lookup.

Where stable identity becomes necessary

Multi-replica, high-change systems

As churn and caller count grow, duplicated client coordination exceeds simple lookup benefits.

Costs Accepted

Declaration

Platform needs explicit service-to-instance membership declaration.

Delayed observation

Eligibility and topology signals are observed over time, not instantaneously.

Propagation windows

Mapping updates are continuous but not atomic across all participants.

Indirection

A platform data-plane component is required even though this chapter does not implement it.

Why is 100 callers x 5 replicas highlighted?

That is 500 partial membership facts reconstructed at the edge of applications, versus one platform-owned service membership view.

Did this chapter solve readiness eligibility policy?

No. It only derives that mapping needs eligibility input. Policy details remain open.

Investigation Exercise (Optional)

Run a cancellable synthesis trace and audit each claim boundary.

Prediction

Predict which ownership model changes when Pod replacement occurs.

Experiment

Run four-step trace from lookup timing to platform-owned stable identity.

Observation

Record where delayed or partial observation changes client behavior.

Reflection

State what was proven as architecture and what remains mechanism or policy.

o o o
Run the synthesis trace after writing your prediction.

Bridge to INV-022

Stable address is earned.
Human-readable name resolution is unresolved.

Service identity can now outlive Pod replacement.

Clients no longer need to track every replica address directly.

Platform owns the changing mapping behind stable service identity.

But applications are configured with names, not raw addresses.

The next mystery asks who turns names into current service identity.

Stable service identity now exists; resolving human-readable names to that identity is still open
Given a service name, who knows the current address now?

Next Investigation

INV-022 - The Name Resolution Problem

Stable address exists. Name-to-address discovery remains unresolved.

Deliberate Simplifications Ledger

Data-plane realization for stable address forwardingOpen backlog
Service access scopes and boundary exposure modelINV-023
Eligibility and readiness determination for traffic inclusionOpen backlog
Traffic distribution policy, headless/ExternalName variants, and balancing algorithmsOpen backlog
Name-to-address resolutionINV-022