Home /Release engineering

Your first incident runbook: decisions a small team should prepare

Prepare ownership, containment, recovery and communication steps before a customer-facing incident forces rushed decisions.

Weekly fieldnotes / 15

Archive week: . This retrospective edition was published in September 2026.

When it breaks, who does what? — Release engineering
CodeSignOff Fieldnotes · Release engineeringDownload banner ↗

Write for the person responding under pressure

A first runbook should be short enough to use during an interruption. Start with how to recognise the incident, who coordinates it and where the team records actions. Include an escalation contact when the usual engineer is unavailable. Avoid a document that lists every hypothetical disaster while leaving basic access unresolved. The responder needs a path from observation to a safe next decision.

Separate coordination from investigation

Even a small team benefits from knowing who is investigating and who is keeping the incident record current. Google’s SRE incident guidance provides a useful model for clear roles. Adapt the scale to your team rather than copying a large organisation’s ceremony. Record material actions and their times so the team can understand what changed. Without that record, simultaneous fixes can make the system harder to diagnose.

Rehearse one realistic scenario

Choose a specific failure, such as paid orders not receiving confirmation. Identify the customer impact, the first check and a reversible containment action. Discuss whether new orders should be paused. Use a tabletop exercise or controlled staging failure rather than disrupting customers. The purpose is to uncover missing permissions, unclear ownership and assumptions about external support before those gaps become part of a live incident.

Define recovery evidence

Agree what must be true before the team declares service restored. A restarted process may still leave incomplete orders or delayed jobs. Check the affected customer journey and identify records needing reconciliation. Assign responsibility for follow-up work and customer communication through the authorised channels. Avoid announcing a complete resolution merely because an infrastructure graph recovered while the business workflow remains inconsistent.

Improve the system after the event

Review the sequence without reducing it to who made a mistake. Identify changes that improve detection, limit consequences or simplify recovery. Keep actions small enough to assign and verify. Update the runbook with what actually helped. A document that survives only as an onboarding attachment is not operational preparation; a rehearsed procedure connected to access, monitoring and ownership can materially improve the next response.

Sources & further reading

Reference material for the guidance and examples above. Where included, community discussions provide context rather than verified incident evidence.

Retrospective weekly fieldnote, prepared with AI assistance and published on 12 September 2026. Examples are illustrative; this is not a client case study or a claim about events in the assigned week.

Continue reading

← Back to all insights
A clear next step

Put these questions to your own codebase.

Tell us what you’re building and what’s coming next.
We’ll help you scope the right review.