Deciding storage must exist
Scheduling, orchestration, and lifecycle decisions belong to the platform alone.
Investigation 027 - Persistent State Principles
Persistence and placement are settled contracts. Both quietly assumed the platform already understood how every storage backend does its work.
Begin the investigation downPrologue
The platform knows that a volume must exist. It does not own the infrastructure capable of creating one. Cloud block devices, storage appliances, and distributed file systems each speak a different language, expose different capabilities, and evolve on their own schedule.
Persistent storage has a platform-owned lifecycle.
The platform can simply learn how each backend creates a volume.
A new storage system arrives. The platform has never seen it.
Can a platform grow without limit by accumulating knowledge of the systems it depends on?
First Principles
The platform decides that storage must exist. A storage system decides how it comes into being. Those are different responsibilities, owned by different organizations, evolving for different reasons.
Scheduling, orchestration, and lifecycle decisions belong to the platform alone.
Cloud block devices, storage arrays, and distributed volumes each implement creation differently.
Responsibilities that change for different reasons belong in different components — including the platform's own boundary with the systems it depends on.
Knowledge accumulates wherever ownership and implementation share one boundary.
Every delegated storage operation crosses a boundary the platform cannot observe directly. Whether it is completing slowly or has failed outright is deferred, not resolved, by this investigation.
Naive Architecture
The platform owns the storage lifecycle. It therefore implements storage directly — one hardcoded driver per backend, written by the platform team, shipped inside the platform itself.
The Architecture That Almost Worked
Applications see one storage API. Operators run one deployment. Every supported backend's implementation lives inside the same platform process and ships with the same platform release.
Vendor-specific interfaces never reach applications directly.
Each driver executes inside the platform's own process and release.
For a small, fixed set of backends owned by one organization, this may be exactly right. The failures that follow test whether that condition still holds.
Breaking Our Design
Each experiment starts from its own clean state and stops at its assigned discovery.
Pressure. A storage vendor ships a product the platform has never seen. Supporting it means writing another driver and shipping another release.
Prediction. Adding one backend seems harmless. Scaling to the size of a real ecosystem may not be.
Add a first backend, then add more, then scale to ecosystem size.
Each backend requires its own driver, its own tests, its own ongoing maintenance.
The platform's codebase and maintenance surface grow with every backend, whether or not its own architecture changed.
Platform knowledge, driver codebases, tests, and maintenance grow linearly with the external storage ecosystem.
What happens when that vendor's code runs inside the platform's own process?
No conclusion yet about shared failure or release schedules.
Pressure. A driver written by a storage vendor contains a defect — a memory leak, a deadlock, a crash.
Prediction. The defect belongs to the vendor. Whether it stays contained depends on where the driver executes.
Embed an independently owned vendor driver, then trigger its defect.
The driver leaks memory, deadlocks, then crashes inside the platform's own process.
Every backend sharing that process degrades along with it.
Execution location expands the platform's failure responsibility; a vendor defect can affect unrelated orchestration and every other backend in the same process.
What happens when platform and vendor releases must move at different speeds?
No conclusion yet about independent release schedules or a final contract.
Pressure. The storage vendor and the platform team ship on their own schedules, for their own customers.
Prediction. Diverging versions will force a coordinated release neither side can complete alone.
Diverge vendor and platform versions, then force a coordinated release.
One side waits on the other; feature and version mismatches appear on both.
Every backend the platform has accepted becomes another coordination requirement per release.
Independent ownership requires independent schedules; a shared implementation and release boundary is unsustainable.
What replaces a shared implementation and release boundary neither side can control alone? Chapter 07 formalizes the answer once all three experiments are complete.
Protocol, transport, runtime isolation, and compatibility negotiation mechanics remain deferred.
The Turning Point
Communication crosses the boundary. Implementation does not.
The Storage Interface Contract
The platform owns a minimal stable contract. Every storage system owns its own implementation. Communication crosses that boundary. Implementation does not.
Ask that a volume come into existence, however the backend provisions it.
Ask that a volume be released when it is no longer needed.
Ask that storage become reachable where computation executes.
Ask that the storage no longer be tied to that execution location.
These four generic operations are not the complete set any real interface exposes. They are the shape this investigation derived — not a specification.
Only Now: Kubernetes
Kubernetes adopts the Container Storage Interface (CSI) as the realization of the boundary this investigation derived — a vendor-neutral standard interface. Out-of-tree CSI drivers can be packaged and released separately from Kubernetes itself, but every driver must still maintain and test compatibility with the specific Kubernetes and CSI versions it supports; independent release does not mean compatibility independence.
The platform calls the contract. The driver translates each call into whatever operations its backend actually requires.
CSI operations realize the four generic lifecycle categories through a stable API.
An out-of-tree CSI driver, owned by the vendor, performs the backend-specific steps.
Protocol and runtime mechanics belong to a later investigation, not this one.
CSI has richer operations and capability negotiation than the four generic categories this investigation derived, including separate Identity, Controller, and Node services — snapshots, expansion, and staging among them. CSI does not make a stalled operation distinguishably slow versus failed, and it does not make a partially completed operation atomic. Both remain open questions.
Engineering Reflection
A platform cannot grow indefinitely by accumulating knowledge of the systems around it. Successful platforms define stable contracts instead. The platform owns the contract; external systems own their implementations.
When the platform and its storage infrastructure are owned by one organization and released together, embedding implementations may be the simplest correct design.
When storage systems are owned by other organizations and must evolve independently, a stable boundary is what makes that possible.
The contract must remain useful across many different storage systems.
Every implementation must satisfy the contract even when its native behavior differs.
New capabilities must extend the interface without breaking existing implementations.
Backend-specific optimizations remain intentionally invisible behind the boundary.
Investigation Exercise
Five storage vendors each release a major new version this year. Predict where engineering effort must go.
List the platform team's responsibilities. Separate orchestration work from storage-implementation work.
Notice which group keeps growing even though the platform's own purpose never changed.
Ask what changes once the platform owns only the contract, not the implementations.
Bridge to INV-028
The platform no longer grows by accumulating knowledge of every storage system it encounters.
It defines a minimal stable contract; storage systems own their implementations independently.
Kubernetes adopts CSI as the realization of that contract — one instance of the boundary this investigation derived.
A failed Pod can be replaced, and the replacement can reattach to the same durable volume.
The data survives. Its name, its network identity, and its membership in a cluster do not automatically survive with it.
This investigation inherits architectural decomposition and independent responsibility from earlier investigations. The Kubernetes CSI design proposal documents the concrete move from compiled-in volume plugins to out-of-tree third-party drivers behind a stable interface. CSI is one realization of the broader boundary principle derived here; this investigation does not claim that CSI invented stable interfaces or organizational boundaries.