Incident Runbook
When the cockpit reports trouble, work the matching play top to bottom. Each play opens at a severity floor, walks an ordered set of checkable steps, and names the condition that hands the incident up the escalation ladder. Severities match the escalation board.
Node offline
A worker node drops out of the cluster heartbeat (missed > 2 intervals).
Confirm the loss, drain in-flight work, and restore quorum before capacity bites.
Enters at Warning- 1
Confirm the node is truly gone — cross-check the cluster board against the uptime feed.
Expect: Heartbeat absent on both surfaces, not a transient blip.
- 2
Cordon the node so the scheduler stops placing new work on it.
Expect: Node shows as unschedulable; pending queue stops growing for it.
- 3
Reschedule in-flight tasks owned by the node onto healthy peers.
Expect: Owned tasks reappear on surviving nodes within one scheduling interval.
- 4
Check remaining headroom — if quorum or capacity is at risk, add a replacement node.
Expect: Cluster reports quorum healthy and headroom above the warning floor.
Service DEGRADED
A service reports DEGRADED: serving traffic but breaching its latency or error budget.
Stabilize the user-facing path first, then find and clear the cause without making it worse.
Enters at Warning- 1
Read the live status board — identify which dependency or budget is breaching.
Expect: A single failing signal (latency, error rate, or saturation) is isolated.
- 2
Shed or throttle non-critical load to protect the core path.
Expect: Error budget burn rate flattens; core requests keep succeeding.
- 3
Correlate against the most recent deploy and the dependency map.
Expect: The change window or upstream that lines up with onset is identified.
- 4
Apply the smallest safe mitigation — roll back the suspect change or fail over the dependency.
Expect: Status returns to healthy and holds for a full observation window.
Escalation hand-off
A play stalls past its escalate-when condition, or severity ratchets to critical.
Hand the incident to a wider circle cleanly, with state preserved and a commander named.
Enters at Critical- 1
Declare the incident and open a tracked thread — stop solo-debugging.
Expect: Incident is recorded with a stable id and the current severity floor.
- 2
Page the next rung per the escalation ladder (owner → area lead → incident commander).
Expect: The right circle acknowledges within the rung window.
- 3
Freeze dependent transitions so the incident cannot widen while hands change.
Expect: Blocked transitions are visible on the board; no new dependent work starts.
- 4
Hand off with a one-paragraph state summary: trigger, steps tried, current severity.
Expect: Commander restates the state back; ownership is unambiguous.
Runbook plays are illustrative cockpit project-tracking policy authored in src/data/runbook, not product data. Severities are reused from src/data/escalation-policy so this board and the escalation ladder stay in sync.