All screens
42 days

Incident Runbook

When the cockpit reports trouble, work the matching play top to bottom. Each play opens at a severity floor, walks an ordered set of checkable steps, and names the condition that hands the incident up the escalation ladder. Severities match the escalation board.

Node offline

A worker node drops out of the cluster heartbeat (missed > 2 intervals).

Confirm the loss, drain in-flight work, and restore quorum before capacity bites.

Enters at Warning
  1. 1

    Confirm the node is truly gone — cross-check the cluster board against the uptime feed.

    Expect: Heartbeat absent on both surfaces, not a transient blip.

  2. 2

    Cordon the node so the scheduler stops placing new work on it.

    Expect: Node shows as unschedulable; pending queue stops growing for it.

  3. 3

    Reschedule in-flight tasks owned by the node onto healthy peers.

    Expect: Owned tasks reappear on surviving nodes within one scheduling interval.

  4. 4

    Check remaining headroom — if quorum or capacity is at risk, add a replacement node.

    Expect: Cluster reports quorum healthy and headroom above the warning floor.

Escalate whenQuorum cannot be restored within the warning window, or a second node drops.

Service DEGRADED

A service reports DEGRADED: serving traffic but breaching its latency or error budget.

Stabilize the user-facing path first, then find and clear the cause without making it worse.

Enters at Warning
  1. 1

    Read the live status board — identify which dependency or budget is breaching.

    Expect: A single failing signal (latency, error rate, or saturation) is isolated.

  2. 2

    Shed or throttle non-critical load to protect the core path.

    Expect: Error budget burn rate flattens; core requests keep succeeding.

  3. 3

    Correlate against the most recent deploy and the dependency map.

    Expect: The change window or upstream that lines up with onset is identified.

  4. 4

    Apply the smallest safe mitigation — roll back the suspect change or fail over the dependency.

    Expect: Status returns to healthy and holds for a full observation window.

Escalate whenMitigation does not recover within the warning window, or DEGRADED crosses into an outage.

Escalation hand-off

A play stalls past its escalate-when condition, or severity ratchets to critical.

Hand the incident to a wider circle cleanly, with state preserved and a commander named.

Enters at Critical
  1. 1

    Declare the incident and open a tracked thread — stop solo-debugging.

    Expect: Incident is recorded with a stable id and the current severity floor.

  2. 2

    Page the next rung per the escalation ladder (owner → area lead → incident commander).

    Expect: The right circle acknowledges within the rung window.

  3. 3

    Freeze dependent transitions so the incident cannot widen while hands change.

    Expect: Blocked transitions are visible on the board; no new dependent work starts.

  4. 4

    Hand off with a one-paragraph state summary: trigger, steps tried, current severity.

    Expect: Commander restates the state back; ownership is unambiguous.

Escalate whenAlready at the top of the ladder — keep the incident at the critical floor until resolved by a human.

Runbook plays are illustrative cockpit project-tracking policy authored in src/data/runbook, not product data. Severities are reused from src/data/escalation-policy so this board and the escalation ladder stay in sync.