Investigation 011 · Failure Detection

A machine stops responding.
Has it actually failed?

One machine can never inspect another machine's processor, memory, or operating system. It can observe only communication — and missing communication is never absolute proof of failure. This is the mystery of failure detection.

Begin the investigation ↓

Author’s Note

We are not searching for a better mechanism. We are searching for certainty itself.

Every other investigation in this book follows the same shape: an obvious solution breaks, a better one replaces it, reality breaks that too, until only one architecture survives.

This investigation is different. Can one machine ever truly know that another machine has failed? As you will discover, that confidence is misplaced — and that single limitation changes everything that follows.

Prologue

The worker who stopped answering the phone

You manage a team spread across cities. Every worker reports by phone. For months, everything runs smoothly.

Then one morning, a worker stops answering. You call again. Nothing. Has she fallen sick? Is her phone dead? Is she stuck in a tunnel? Has she disappeared entirely?

From where you stand, every possibility looks the same. All you know is that communication has stopped. Waiting forever isn’t an option — but acting too quickly risks assigning her work to someone else while she is still doing it.

The mystery. How can one machine ever know whether another machine has actually failed?

First Principles

Machines never observe failure directly. They observe only communication.

In everyday life, certainty often comes from direct observation — we can inspect a light bulb, open a car's hood, look at a frozen screen. Distributed systems don’t have that luxury.

Each machine exists in its own world. It cannot see another machine’s processor, inspect its memory, or check whether its operating system is running. The only thing it can observe is communication. Messages arrive, or they don’t.

Suppose a machine sends a message and receives no reply. What has it actually learned? Only one fact: communication did not occur as expected. Not why.

Only one fact: communication did not occur as expected. Not why.
First architectural principle. Machines never observe failure directly. They observe only communication. Everything in this investigation follows from that single idea.

The Impossible Goal

Perfect failure detection: always correct, never late, never a false alarm

The requirements sound obvious. A healthy machine should never be mistaken for a failed one. A failed machine should never be mistaken for a healthy one. Detection should be immediate. Recovery should happen without delay.

It sounds like an engineering problem waiting for a clever solution. The obvious first attempt: just ask. If it replies, it’s alive. If not, it must have failed.

The Architecture That Almost Worked

Just ask

Machine A sends Machine B a message: “Are you still there?” If B replies, everything is normal. If not, A concludes B has failed. No shared state. No coordination. Just a question and an answer.

Machine A: "Are you still there?"
→
Machine B

For a moment, it feels solved. When a machine fails, it stops responding — the absence of a reply becomes the signal.

The uncomfortable question. What if B never receives the message? Or receives it, but the reply is lost? Congested network, overloaded switch, a disconnected sender — in every case, A observes exactly the same thing: silence.

The architecture assumed silence has only one explanation. Reality offers many. The problem is no longer whether B is alive — it’s whether silence can ever tell us why communication stopped.

The problem is no longer whether B is alive — it’s whether silence can ever tell us why communication stopped.

Lab — What Silence Actually Means

Machine A asks. Machine B stays silent. Reveal what really happened behind that identical silence.

Machine A observes: silence. Cause unknown.

Breaking Our Design

Every design fails for the same reason

Silence never explains itself. It tells us that communication failed — never why. We deliberately try every reasonable fix, and each one fails the same way.

Episode 1 — Silence Is Ambiguous

Machine A sends a message. No reply arrives. Has B failed? Perhaps — but the message may still be traveling, a router may have dropped it, B may be overloaded, or A itself may be disconnected.

Every explanation produces identical evidence: no reply. If we immediately declare every silent machine dead, healthy machines get treated as failures. If we never declare them dead, real failures go unhandled.

Episode 2 — Heartbeats Change Nothing

Instead of asking once, every machine announces “I’m still here” at regular intervals. It feels like a major improvement — a continuous stream instead of one request.

But if a heartbeat never arrives, we’ve learned nothing new — only another missing message. Crash, dropped packet, congestion, paused process: all identical. Heartbeats give more opportunities to observe. They don’t create certainty.

Episode 3 — The Timeout Dilemma

Perhaps we were just too impatient — wait longer before deciding. But how long? Short timeouts detect failures fast but mistake temporary delays for permanent failure. Long timeouts reduce false alarms but delay real recovery.

Discovery: no timeout can maximize both speed and confidence. Choosing one isn’t about finding the "correct" value — it’s deciding which mistake is less harmful.

Episode 4 — No One Sees the Whole Network

Maybe one observer just isn’t enough. Let three machines watch the same participant. One sees perfect health. One experiences delays and suspects failure. One loses connectivity entirely and sees total disappearance.

All three are correct — each reports exactly what it can see. There is no single vantage point from which the whole system is observable. More observers add information. They don’t create certainty where none can exist.

No timeout can maximize both speed and confidence.
Compact review

Lab — Choose Your Timeout

Slide the timeout and watch the tradeoff between false evictions and slow recovery play out.

Timeout: 15sBalanced

Move the slider to see how the same missing heartbeat gets judged differently.

Lab — Three Observers, Three Realities

The same machine, watched by three observers on three different network paths.

No observer is wrong. Each is reporting exactly what reached it.

Turning Point

Perfect failure detection does not exist

We tried direct requests. Continuous heartbeats. Longer waits. More observers. Every design failed for the same reason: no machine can directly observe another machine’s state. It can observe only communication — and communication is never a perfect reflection of reality.

The goal was never achievable as stated. The real question was never “how do we detect failure with certainty?” It was always “how do we make safe decisions despite the absence of certainty?”

“How do we make safe decisions despite the absence of certainty?”

The Architectural Contract

Six rules for deciding under uncertainty

1. Observe communication, not failure

Base every decision on communication received — or not received.

2. Suspicion, not proof

A missing message is evidence something may be wrong. Never definitive proof.

3. Accept the speed/accuracy tradeoff

Faster detection risks false suspicion. Slower detection delays recovery.

4. Expect observers to disagree

Different observers see different networks. Disagreement is normal, not a flaw.

5. Allow suspicion to be reversed

A suspected machine may resume communication. Suspicion is not permanent failure.

6. Preserve correctness despite imperfect detection

The system must remain correct even when suspicions are occasionally wrong.

The reframe. These principles do not eliminate uncertainty — they acknowledge it. Reliable distributed systems are not built by discovering perfect failure detectors. They are built by accepting that perfect failure detectors cannot exist.

Engineering Reflection

Kubernetes never claims to know a node has failed

Each node runs a kubelet, which periodically communicates with the control plane. Modern Kubernetes uses Node Lease objects for these frequent updates, while the Node object itself updates less often. Every successful renewal is evidence of continued communication — not proof of health.

If the control plane stops receiving lease renewals within the configured grace period, the NodeLifecycleController marks the node as NotReady. Notice the wording: it declares only that expected communication has stopped, not that the node crashed.

Kubernetes applies taints like node.kubernetes.io/not-ready and node.kubernetes.io/unreachable to influence scheduling and eviction — while still allowing recovery if communication resumes.

Costs Accepted

Detection latency is unavoidable. A failure isn’t detected until the timeout expires — a direct consequence of the safety guarantee.

False evictions are a real cost. A partitioned-but-alive node may have its work evicted and rescheduled elsewhere; when the partition heals, duplicate instances may briefly exist. Versioned truth and leadership contracts from earlier investigations absorb that risk.

Heartbeat write pressure at scale. Every node writes lease renewals to a shared store at regular intervals — background write load on etcd proportional to node count, not workload.

Timeout tuning is never final. Grace periods must be recalibrated as cluster conditions change — a value correct today may be wrong tomorrow.

These costs are accepted because the alternative is worse: an architecture that waits for certainty before acting would wait forever.

Investigation Exercise

Hypothesis. Failure detection is not the act of proving another machine has failed — it is the act of making decisions based on missing communication.

Experiment. Run two machines exchanging a heartbeat every second. Introduce, one at a time: stopping the sender, disconnecting the network, delaying traffic, pausing the sender. Record what the receiver actually observes in each case.

Reflection. The receiver never observes the cause — only that expected communication stopped. Different failures often look identical. That's why production systems rely on heartbeats, leases, grace periods, and recovery policies instead of attempting to prove failure.

Bridge to INV·012

Detecting absence is not the same as resolving what was left behind

We’ve reached an uncomfortable conclusion: a distributed system can never know with certainty whether another machine has failed. It can only observe communication and act on incomplete information. That uncertainty cannot be eliminated — only managed.

But a new mystery emerges. Suppose a machine really does disappear. It wasn’t working alone — it owned objects, created resources, established relationships. Those things don’t vanish just because their creator is gone.

Failure detection tells us that an owner may no longer be participating. It says nothing about the fate of what that owner owned.

Should those resources remain forever? Be deleted immediately? Who decides? That mystery — how systems preserve correct ownership throughout the entire lifecycle of their resources — is the subject of the next investigation.

A silent machine becomes suspected rather than certainly failed, while its objects, resources, and relationships remain with their lifecycle unresolved

Deliberate Simplifications Ledger

We deliberately postponedOwned by
What happens to objects owned by a machine that stops communicatingINV·012 — Owner References
How Kubernetes decides when to evict Pods from a NotReady node (taint-based eviction, grace periods)Implementation detail — NodeLifecycleController policy
The FLP impossibility result — formal proof that no algorithm guarantees both safety and liveness with even one faulty processMathematical foundation — Fischer, Lynch, Paterson (1985)
How network partitions force a system to choose between consistency and availabilityFuture investigation — CAP theorem territory
Clock skew between nodes and its effect on lease expiry accuracyImplementation detail — NTP and monotonic clock considerations

Sources

Research: Impossibility of Distributed Consensus with One Faulty Process — Fischer, Lynch, Paterson, JACM 1985; Unreliable Failure Detectors for Reliable Distributed Systems — Chandra & Toueg, JACM 1996; Leslie Lamport, Time, Clocks, and the Ordering of Events in a Distributed System, CACM 1978.

Official Documentation: Kubernetes Node Lease design (kubernetes/enhancements/issues/589); NodeLifecycleController node condition, taint, and eviction logic (pkg/controller/nodelifecycle); kubelet --node-status-update-frequency and kube-controller-manager --node-monitor-grace-period flags.

Books: Designing Data-Intensive Applications — Martin Kleppmann (Chapter 8); Programming Kubernetes — Hausenblas & Schimanski.

Next: INV·012 — Owner References