Investigation 022 - Cluster Networking Principles

The name stayed.
The address changed.

INV-021 gave each service a stable platform-owned address. Applications still use names. When that address changes, who keeps the name connected to current platform state?

Begin the investigation down
postgres10.96.14.27then 10.96.18.42

Author's Note

Keep service identity separate from current location.

This chapter inherits stable service addresses from INV-021. It does not rediscover them.

It asks how applications can retain stable names while the platform changes the addresses those names represent. We begin with a table, let it work, and stop at an architecture-level resolver contract.

Applications know what they need. The platform knows where it currently is.

Prologue

The lookup succeeds. The connection goes nowhere.

Orders has always called postgres. Yesterday, an operator mapped that name to 10.96.14.27. Today, the platform replaces the service address with 10.96.18.42.

Application intent

Connect to the same logical service: postgres.

Platform state

The service now owns 10.96.18.42.

Copied mapping

The static entry still says 10.96.14.27.

Result

A valid-looking answer describes yesterday's platform.

First Principles

Identity and location have different lifetimes.

A name expresses durable application intent. An address expresses current platform location. Resolution connects them without pretending they change together.

Identity

Stable name

Describes which service the application intends to reach.

Location

Current address

Describes where the platform currently receives traffic for that service.

Evidence

Observed answer

A derived observation may lag authoritative state; it is not instantaneous global truth.

Observation is not authority. A resolver answer can be useful while still requiring convergence after change.

Naive Architecture

Maintain one simple table.

Operators map each service name to its stable address. Applications read the table and avoid embedding infrastructure addresses directly.

Applicationpostgres
Static Mappingoperator maintained
10.96.14.27service address
Benefit

No running resolver

The design adds no continuously operating component.

Benefit

Names stay stable

Applications remain free of direct address configuration.

Honesty

Valid at small scale

Few fixed services with rare changes can use this design well.

The Architecture That Almost Worked

Maintain one mapping. Distribute it everywhere.

Operators update the shared table whenever service state changes, then send that copy to every application group that needs it.

Service Namepostgres.default
Shared Tablecopied to applications
Service10.96.14.27
The platform owns authoritative service state. The table has quietly become an independently maintained copy.

Breaking Our Design

Four independent episodes remove the second source of truth.

EPISODE 01

Static Entries Go Stale

Pressure. Platform replaces the postgres service address while a copied mapping remains unchanged.

Prediction. A syntactically valid mapping should still provide a trustworthy answer.

Experiment

Query, replace address, query untouched copy.

Observation

Lookup succeeds with the old valid-looking address.

Failure

The copy silently diverges from authority.

Discovery

A copy cannot detect divergence from itself.

Next pressure

Can updates repair the design?

Boundary

Detection is not solved here.

Independent simulation: copied state versus platform state.

Platformpostgres -> 10.96.14.27
Static copypostgres -> 10.96.14.27
Lookupnot queried

Prediction: a valid entry should remain trustworthy.

EPISODE 02

New Services Require Coordination

Pressure. Notifications becomes ready at 10.96.33.18, but application groups hold separate mapping copies.

Prediction. Platform readiness should imply that applications can discover the service.

Experiment

Deploy once, then distribute independently.

Observation

Ready service remains undiscoverable to stale groups.

Failure

Two operations lack shared completion.

Discovery

Readiness does not imply discoverability.

Next pressure

What does distribution cost at scale?

Boundary

No automation is introduced yet.

Independent simulation: deployment and mapping distribution.

Platformnotifications absent
Application groups0 of 3 can discover it
Completionno operation started

Prediction: platform readiness should imply discoverability.

EPISODE 03

The Update Problem at Scale

Pressure. Two hundred services are copied into an increasing number of application instances.

Prediction. One service change should remain one small mapping operation.

Experiment

Scale instances, then change one service.

Observation

Facts and update targets grow with copies.

Failure

One change becomes participant-wide work.

Discovery

Cost grows with copies and participants.

Next pressure

Can applications stop owning copies?

Boundary

No universal threshold is claimed.

Independent simulation: 200 services across application instances.

Instances10
Mapping facts2,000
Change targetsnot run

Prediction: one service change should remain one small operation.

EPISODE 04

The Dynamic Resolver Requirement

Pressure. Applications need current answers without owning service state or independently distributed copies.

Prediction. A platform-derived view can converge after authoritative change while remaining honest about observation lag.

Experiment

Query, change authority, query before refresh, reconcile, query again.

Observation

Derived observation can lag, then converge.

Success

Current answer returns after explicit refresh.

Discovery

Platform owns authority; resolver owns derived view.

Contract

Continuously reconcile resolution.

Boundary

No instant propagation or global truth.

Independent simulation: authoritative state and resolver-owned derived view.

Authoritative service statepostgres -> 10.96.14.27
Resolver-derived viewpostgres -> 10.96.14.27
Query answernot queried

Prediction: a derived view may lag, then converge explicitly.

Optional Episode Review

The Turning Point

Do not distribute copies of platform knowledge.
Derive the answer where that knowledge is owned.

The platform remains authoritative for service state. A resolver owns only the continuously reconciled view used to answer names.

The Name Resolution Contract

Five responsibilities. No required protocol.

Whenever names outlive addresses, resolution must continuously derive from the authoritative state that defines current location.

Contract 1

Stable names

Applications identify services by logical name, not infrastructure address.

Contract 2

Platform ownership

The platform owns address resolution; applications do not maintain mappings.

Contract 3

Derived resolution

Answers converge from authoritative service state. The resolver owns a derived view, never service truth.

Contract 4

One consistent view

Applications receive answers derived from the same authoritative platform state rather than divergent copies.

Contract 5

Invisible address churn

Service address changes do not require application configuration changes.

The contract does not require DNS, record formats, caching rules, query transport, or instant propagation.

Only Now: Kubernetes

Kubernetes cluster DNS commonly realizes this contract.

An application can continue using postgres.default.svc.cluster.local while the platform-owned Service address changes. The observed answer follows the cluster's current Service description through a derived view.

CoreDNS is a common implementation of Kubernetes cluster DNS. It is not the universal architecture, and this investigation does not enter its internals.

The architecture is continuously derived name resolution. DNS and CoreDNS are realizations of that contract.

Engineering Reflection

Timeless Engineering Principle

When identity outlives location, keep the name stable and continuously derive its current location from authority.

Architectural Honesty

Static mappings remain valid

Small, low-change systems

A handful of services, rare address changes, and one visible distribution path may justify the simpler design.

Derived resolution becomes necessary

Dynamic, independently scaled systems

Frequent churn and many application instances make copied mappings a coordination burden.

Costs Accepted

Availability

Applications now depend on a resolver capability.

Observation

The resolver must observe authoritative service state.

Propagation

Derived answers converge through non-zero delay.

Indirection

Every lookup crosses additional platform-owned infrastructure.

Did we prove a numerical point where static mappings fail?

No. The experiment exposes how cost grows with participants. The acceptable threshold is policy and context, not universal architecture.

Does one resolver answer prove current global truth?

No. It is a derived observation. Correctness requires explicit convergence from authoritative state, not a claim of instantaneous propagation.

Investigation Exercise

Audit one address change without extending the contract.

Prediction

Predict what each application observes when only half receive a changed mapping.

Experiment

Trace one authoritative address change across one hundred copied mappings.

Observation

Separate authoritative state, stale evidence, and current derived evidence.

Reflection

Identify who should own service state and who should own its derived view.

o o o
Run the synthesis trace after writing your prediction.

Bridge to INV-023

Internal applications can now find services.
External requests still cannot enter.

Service names remain stable while addresses change.

Applications no longer maintain copies of platform knowledge.

The resolver continuously derives a view from authoritative service state.

But a browser knows nothing about internal names or virtual addresses.

A public request arrives at the platform boundary asking for entry.

Internal name resolution reaches a service while an external request stops at the unresolved platform boundary
Name resolution discovers where a service exists inside the platform. External access asks how a request enters it.

Next Investigation

INV-023 - The External Access Problem

Which internal service should receive this external request?

Intellectual Lineage

This investigation inherits the feedback-loop and reconciliation model from control theory, together with platform precedent from Borg and Omega. It applies that inherited lineage to a resolver-owned derived view; it does not claim that Borg or Omega introduced cluster DNS.

Deliberate Simplifications Ledger

CoreDNS internals, plugins, record types, and query transportOpen backlog
Caching, TTLs, negative caching, and Pod lookup configurationOpen backlog
Headless Services and Pod-level namingOpen backlog
External traffic entryINV-023
External DNS, multi-cluster resolution, forwarding, and DNS securityOpen backlog