Routes target machines
Routers are built to forward toward host and subnet destinations.
Investigation 020 - Cluster Networking Principles
We inherit two contracts without rediscovery: INV-018 established that a running Pod instance owns its own network identity, and INV-019 established that a node must own finite resource boundaries. The mystery now is cross-node communication: why does a packet addressed to a remote Pod vanish when physical links are healthy?
Begin the investigation downPrologue
Orders runs on one node. Inventory runs on another. Orders sends to Inventory's Pod address. The physical network carries packets between machines, but it has never been taught where Pod addresses live.
Orders Pod sends traffic.
Routers evaluate destination against machine routes.
No route exists for the destination Pod address.
The machines are linked; workload reachability is still unresolved.
First Principles
Physical networks route machine addresses. Workload identity answers which running instance a packet is for. These responsibilities overlap on one node, but diverge immediately across nodes.
Routers are built to forward toward host and subnet destinations.
A Pod IP identifies one running instance, not the node itself.
A workload address can be correct locally but meaningless to the physical fabric.
Do not reveal the answer early: first isolate each failure caused by forcing workload communication through machine identity.
Naive Architecture
To avoid creating a new network model, every Pod is addressed through its node IP plus a selected port.
The Architecture That Almost Worked
If one Pod runs per node, or if operators coordinate unique node ports manually, traffic appears stable and the design seems successful.
No local port contention exists.
Teams avoid collisions by agreement and sequencing.
The design has not proven independent workload addressability.
The design works while reality is cooperative. Breaking episodes test when cooperation ends.
Breaking Our Design
Pressure. Two independent Pods on one node both need tcp/8080.
Prediction. If borrowed node identity is sufficient, both should bind independently.
Both Pods bind node IP port 8080.
First bind succeeds; second bind is rejected.
Port ownership becomes a shared node bottleneck.
Shared node identity breaks independent Pod execution.
Try translation to keep one node identity while avoiding collisions.
This is an endpoint ownership conflict, not an app bug.
Node identity cannot safely represent many independent listeners.
Run the bind attempt to test shared node endpoint ownership.
Pressure. Keep one node IP while allowing both Pods to keep internal port 8080.
Prediction. Distinct node ports translated to each Pod should restore delivery.
Map node :30001 and :30002 to separate Pod :8080 listeners.
Both requests arrive at their targets.
Connectivity appears repaired.
Translation can hide port collisions.
Now inspect what receiver believes the sender identity is.
Allow this repair to appear successful before deeper tests.
Delivery restored does not yet prove identity correctness.
Configure translation and send both requests.
Pressure. Destination-port translation restored delivery. Now add source masquerading on the forwarded path while the receiver relies on source identity.
Prediction. If this additional translation preserves identity, the receiver should still see the caller Pod IP.
Enable source masquerading, then inspect the receiver log.
Receiver records node IP as sender.
End-to-end source identity is obscured.
Connectivity and identity are different guarantees.
Try direct Pod-addressed routing without translation.
This is a separate identity test, independent from Episode 2 setup.
A node-forwarded identity is not workload identity.
Enable source masquerading and inspect receiver-side identity.
Pressure. Preserve Pod identity and address destination Pod directly across nodes.
Prediction. If physical routing already knows Pod addresses, packet should arrive.
Send direct packet to remote Pod address 10.42.7.3.
Physical route lookup returns no destination path.
Route knowledge for Pod space is absent.
Identity correctness alone does not create reachability.
Derive minimum network guarantees workloads require.
This test isolates routing absence without adding a mechanism.
A second network meaning must be maintained by the platform.
Perform direct route lookup for the destination Pod address.
Pressure. Previous failures must be resolved by explicit workload-facing guarantees.
Prediction. A contract set that guarantees identity and direct reachability should close all prior pressures.
Assemble required guarantees one by one.
Missing guarantees reproduce earlier failures.
Complete set resolves causal chain at contract level.
Cluster-wide direct Pod reachability is required.
Move to stable address identity in INV-021.
Do not implement CNI, BGP, overlays, underlays, or plugins here.
Contract earned before mechanism naming.
Assemble all guarantees to complete the causal chain.
Optional Episode Review
The Turning Point
The requirement is now earned as architecture: any Pod must be directly reachable by Pod identity across node boundaries without translation in the path.
Generic Cluster Network Contract
Every workload may assume that Pod identity is unique cluster-wide, directly reachable without address translation, source-preserving end to end, and independent of node placement.
One Pod IP belongs to one running temporary Pod instance at a time.
Any Pod can address any other Pod directly with no translation hop in the path.
Receivers observe true sender workload identity, not a forwarding node identity.
Applications address workloads without learning which node currently hosts them.
Workloads depend on the cluster reachability guarantees, not on the mechanism used beneath them.
Precision boundary: a Pod IP does not survive replacement. Replacement may receive a different Pod IP.
Only Now: Kubernetes
Kubernetes requires Pod-level addressing that remains meaningful across nodes. Workloads communicate by Pod identity, not by coordinating node ports.
This chapter intentionally does not explain IP allocation internals, plugin internals, overlay versus underlay design, BGP, eBPF, or route convergence implementation details.
Name after discovery, not before discovery.
Engineering Reflection
When workload instances are placed and replaced across machines, network assumptions must continue to name workloads directly, not machines that host them.
Shared node identity can remain acceptable for tightly controlled ingress or single-workload-per-node situations.
Independent teams, same-port workloads, and identity-sensitive policies expose the hidden coupling quickly.
Every node must carry enough route knowledge to reach remote Pods.
Route knowledge can be delayed as Pod instances appear, disappear, and are replaced.
The platform must maintain a shared logical map above physical links.
Applications gain simplicity because platform networking absorbs complexity.
Investigation Exercise (Optional)
Predict which fixes only delivery and which preserves identity.
Run the five-step trace from collision to contract assembly.
Separate temporary success, identity corruption, and routing absence.
State the minimum guarantees without naming implementation mechanism.
Bridge to INV-021
A Pod address is now reachable across nodes, but the address belongs to one temporary Pod instance. If that Pod is replaced, the replacement may receive a new address and old callers still target the vanished instance address.
The network can deliver to a Pod address. It cannot make one Pod address survive Pod replacement.