Investigation 027 - Persistent State Principles

The platform knows storage must exist.
Must it also know how?

Persistence and placement are settled contracts. Both quietly assumed the platform already understood how every storage backend does its work.

Begin the investigation down
Cloud block
platform driver
Storage array
platform driver
Distributed volume
platform driver
Next backend
unknown

Author's Note

Persistence and placement are settled. How is not.

INV-025 gave storage a lifecycle the platform owns independently of any process. INV-026 showed that placement of computation and storage must be co-decided under shared topology. We inherit both contracts rather than rediscovering them.

Every time either contract said "the platform provisions a volume," it assumed the platform already knew how. INV-003 established that responsibilities which change for different reasons belong in different components. This investigation asks whether owning storage orchestration also requires owning every storage implementation.

A delegated storage operation crosses a boundary the platform cannot observe directly. Whether it is completing slowly or has already failed remains an open question this investigation defers rather than resolves.

Communication with storage systems is unavoidable. Whether ownership must cross that boundary is the mystery.

Prologue

Someone still has to create the storage.

The platform knows that a volume must exist. It does not own the infrastructure capable of creating one. Cloud block devices, storage appliances, and distributed file systems each speak a different language, expose different capabilities, and evolve on their own schedule.

Foundation

Persistent storage has a platform-owned lifecycle.

Assumption

The platform can simply learn how each backend creates a volume.

Incident

A new storage system arrives. The platform has never seen it.

Mystery

Can a platform grow without limit by accumulating knowledge of the systems it depends on?

First Principles

Orchestration and implementation change for different reasons.

The platform decides that storage must exist. A storage system decides how it comes into being. Those are different responsibilities, owned by different organizations, evolving for different reasons.

What the platform owns

Deciding storage must exist

Scheduling, orchestration, and lifecycle decisions belong to the platform alone.

What it delegates

How storage is created

Cloud block devices, storage arrays, and distributed volumes each implement creation differently.

Bounded responsibility

INV-003, applied to a new boundary

Responsibilities that change for different reasons belong in different components — including the platform's own boundary with the systems it depends on.

Knowledge accumulates wherever ownership and implementation share one boundary.

Every delegated storage operation crosses a boundary the platform cannot observe directly. Whether it is completing slowly or has failed outright is deferred, not resolved, by this investigation.

Naive Architecture

One driver. One backend. One platform.

The platform owns the storage lifecycle. It therefore implements storage directly — one hardcoded driver per backend, written by the platform team, shipped inside the platform itself.

New Backendvendor infrastructure
Platform Team Writes Driverone implementation
Platform Releasestorage support ships

The Architecture That Almost Worked

One uniform API. Every backend embedded inside it.

Applications see one storage API. Operators run one deployment. Every supported backend's implementation lives inside the same platform process and ships with the same platform release.

What users see

A single, uniform API

Vendor-specific interfaces never reach applications directly.

What the platform contains

Every backend implementation

Each driver executes inside the platform's own process and release.

For a small, fixed set of backends owned by one organization, this may be exactly right. The failures that follow test whether that condition still holds.

Breaking Our Design

Three independent pressures test unlimited growth.

Each experiment starts from its own clean state and stops at its assigned discovery.

EPISODE 01

Every New Backend Changes the Platform

Pressure. A storage vendor ships a product the platform has never seen. Supporting it means writing another driver and shipping another release.

Prediction. Adding one backend seems harmless. Scaling to the size of a real ecosystem may not be.

Experiment

Add a first backend, then add more, then scale to ecosystem size.

Observation

Each backend requires its own driver, its own tests, its own ongoing maintenance.

Failure

The platform's codebase and maintenance surface grow with every backend, whether or not its own architecture changed.

Discovery

Platform knowledge, driver codebases, tests, and maintenance grow linearly with the external storage ecosystem.

Next pressure

What happens when that vendor's code runs inside the platform's own process?

Boundary

No conclusion yet about shared failure or release schedules.

Hardcode one driver per backend and scale the ecosystem.

Backends known0
Driver codebases0
Test suites0
Maintenance owners0
Ecosystem size20+ backends exist industry-wide
Prediction: adding one backend seems harmless. Add more and see whether that holds.
EPISODE 02

Vendor Code Inside the Platform

Pressure. A driver written by a storage vendor contains a defect — a memory leak, a deadlock, a crash.

Prediction. The defect belongs to the vendor. Whether it stays contained depends on where the driver executes.

Experiment

Embed an independently owned vendor driver, then trigger its defect.

Observation

The driver leaks memory, deadlocks, then crashes inside the platform's own process.

Failure

Every backend sharing that process degrades along with it.

Discovery

Execution location expands the platform's failure responsibility; a vendor defect can affect unrelated orchestration and every other backend in the same process.

Next pressure

What happens when platform and vendor releases must move at different speeds?

Boundary

No conclusion yet about independent release schedules or a final contract.

Run vendor-owned code inside the platform's own process.

Vendor drivernot yet embedded
Process boundarynone shared
Driver defectnone triggered
Orchestrationhealthy
Unrelated backendsunaffected
Prediction: the defect belongs to the vendor. Embed the driver to see where it actually runs.
EPISODE 03

The Versioning Deadlock

Pressure. The storage vendor and the platform team ship on their own schedules, for their own customers.

Prediction. Diverging versions will force a coordinated release neither side can complete alone.

Experiment

Diverge vendor and platform versions, then force a coordinated release.

Observation

One side waits on the other; feature and version mismatches appear on both.

Failure

Every backend the platform has accepted becomes another coordination requirement per release.

Discovery

Independent ownership requires independent schedules; a shared implementation and release boundary is unsustainable.

Next question

What replaces a shared implementation and release boundary neither side can control alone? Chapter 07 formalizes the answer once all three experiments are complete.

Boundary

Protocol, transport, runtime isolation, and compatibility negotiation mechanics remain deferred.

Diverge two independently owned release schedules.

Platform versionv1
Vendor driver versionv1
Coordinated releasenot attempted
Wait / mismatchnone observed
Episode outcomenot yet completeived
Prediction: independent versions will eventually force a coordinated release neither side can finish alone.
Review the three experiments
  1. Platform knowledge, driver codebases, tests, and maintenance grow linearly with an external ecosystem that has no natural ceiling.
  2. Execution location, not storage, decides how far a vendor's defect can reach.
  3. Independently owned systems cannot share one release schedule — what replaces that shared boundary is formalized next, in the contract.

The Turning Point

Ask the same four questions.
Stop asking how they're answered.

Platform Knowledgewhat grows without limit
Minimal Stable Contractplatform owned
Backend Implementationstorage owned
Communication crosses the boundary. Implementation does not.

The Storage Interface Contract

One boundary. Four generic operations.

Contract not yet earned. Complete all three experiments before formalizing the contract.

Only Now: Kubernetes

CSI is one realization of the derived contract.

Kubernetes adopts the Container Storage Interface (CSI) as the realization of the boundary this investigation derived — a vendor-neutral standard interface. Out-of-tree CSI drivers can be packaged and released separately from Kubernetes itself, but every driver must still maintain and test compatibility with the specific Kubernetes and CSI versions it supports; independent release does not mean compatibility independence.

The platform calls the contract. The driver translates each call into whatever operations its backend actually requires.

Platform

Calls the contract

CSI operations realize the four generic lifecycle categories through a stable API.

Driver

Translates to the backend

An out-of-tree CSI driver, owned by the vendor, performs the backend-specific steps.

Boundary

What stays deferred

Protocol and runtime mechanics belong to a later investigation, not this one.

CSI has richer operations and capability negotiation than the four generic categories this investigation derived, including separate Identity, Controller, and Node services — snapshots, expansion, and staging among them. CSI does not make a stalled operation distinguishably slow versus failed, and it does not make a partially completed operation atomic. Both remain open questions.

Engineering Reflection

Timeless Engineering Principle

A platform cannot grow indefinitely by accumulating knowledge of the systems around it. Successful platforms define stable contracts instead. The platform owns the contract; external systems own their implementations.

Architectural Honesty

Keep direct integration

When the platform and its storage infrastructure are owned by one organization and released together, embedding implementations may be the simplest correct design.

Introduce the contract

When storage systems are owned by other organizations and must evolve independently, a stable boundary is what makes that possible.

Costs Accepted

Interface stability

The contract must remain useful across many different storage systems.

Conformance

Every implementation must satisfy the contract even when its native behavior differs.

Careful evolution

New capabilities must extend the interface without breaking existing implementations.

Hidden optimization

Backend-specific optimizations remain intentionally invisible behind the boundary.

Investigation Exercise

Trace five vendor releases against a hardcoded platform.

Prediction

Five storage vendors each release a major new version this year. Predict where engineering effort must go.

Experiment

List the platform team's responsibilities. Separate orchestration work from storage-implementation work.

Observation

Notice which group keeps growing even though the platform's own purpose never changed.

Reflection

Ask what changes once the platform owns only the contract, not the implementations.

o o o
Run the synthesis trace after writing your prediction.

Bridge to INV-028

The contract is stable.
But a replacement is not the same member.

The platform no longer grows by accumulating knowledge of every storage system it encounters.

It defines a minimal stable contract; storage systems own their implementations independently.

Kubernetes adopts CSI as the realization of that contract — one instance of the boundary this investigation derived.

A failed Pod can be replaced, and the replacement can reattach to the same durable volume.

The data survives. Its name, its network identity, and its membership in a cluster do not automatically survive with it.

Storage implementations converge on a stable contract before an unresolved identity boundary

Next Investigation

INV-028 - The Stable Identity Problem

Storage preserves memory. Who returns to claim it?

Intellectual Lineage

This investigation inherits architectural decomposition and independent responsibility from earlier investigations. The Kubernetes CSI design proposal documents the concrete move from compiled-in volume plugins to out-of-tree third-party drivers behind a stable interface. CSI is one realization of the broader boundary principle derived here; this investigation does not claim that CSI invented stable interfaces or organizational boundaries.

Deliberate Simplifications Ledger

Storage interface protocols and transportsFuture: Storage Interface Protocol (unassigned)
Runtime isolation of storage implementationsFuture: Storage Driver Isolation (unassigned)
Detecting slow versus failed storage operationsFuture: Storage Failure Detection (unassigned)
Recovery from partially completed storage operationsFuture: Storage Operation Recovery (unassigned)
Interface evolution and backward compatibilityFuture: Storage Interface Evolution (unassigned)