Write an incident runbook

Turn a known failure mode into fast diagnosis, contained mitigation, and clear escalation.

Ready to paste

Replace the bracketed context, then use it in your agent.

Act as a site reliability engineer writing an executable runbook.

An operational failure is important enough that responders should not rediscover the procedure under pressure.

Context to use:
- Failure mode: [describe]
- System context: [reference]

Process:
1. Define the symptom and alerts that activate the runbook.
2. List immediate impact and preflight checks.
3. Provide ordered diagnostic commands with expected interpretations.
4. Describe the least risky mitigation, verification, and rollback.
5. Define escalation, evidence capture, and follow-up ownership.

Constraints:
- Do not include destructive commands without explicit safeguards.
- Do not depend on undocumented tribal knowledge.
- Keep credentials and sensitive values out of the document.

Return:
- Trigger and impact
- Diagnosis steps
- Mitigation
- Verification/rollback
- Escalation and follow-up

Use when

An operational failure is important enough that responders should not rediscover the procedure under pressure.

Expected return

  • Trigger and impact
  • Diagnosis steps
  • Mitigation
  • Verification/rollback
  • Escalation and follow-up

Related prompts

All prompts