Sendabrief logoSendabrief
Free template

Incident response and post-mortem

During an incident nobody has the attention to invent a process, which is exactly why it has to exist beforehand. The part most teams skip is the last third: the post-mortem that turns the fix into a documented change. Without it the same incident recurs and the second occurrence is more expensive than the first.

Download as Word (.docx) 12 stepsNo email, no signup. 12 steps, editable in Word or Google Docs.
Trigger

When any team member observes a customer-affecting failure, or an automated alert fires for a service degradation.

Roles involved
Incident LeadResponderCommunications Owner
Review cadence

After every high-severity incident, and quarterly regardless. This procedure is the one most likely to have been improved by recent experience.

  1. 1

    The person who noticed the issue declares an incident in the agreed channel, stating what is broken and what they observed.

    Responder

    Declaring early and being wrong is much cheaper than deliberating. Nobody should hesitate to declare because they are unsure it qualifies.

  2. 2

    Incident Lead takes ownership by name and states so explicitly in the channel.

    Incident Lead

    One named owner. An incident with no explicit owner has several people assuming someone else is coordinating.

  3. 3

    Incident Lead assesses severity based on customer impact.

    Incident Lead

    Assess by impact, not by how technically alarming the cause looks. A noisy internal failure with no customer effect is not a high-severity incident.

    Customers affected or data at risk: go to step 4

    No customer impact: go to step 6

  4. 4

    Communications Owner notifies affected customers with what is known, what is being done, and when the next update will come.

    Communications Owner

    Timing: Within 30 minutes of severity assessment

    Commit to a next-update time and meet it even when there is no progress. Silence during an incident is what damages trust, not the incident.

  5. 5

    Communications Owner posts updates at the promised interval until resolution.

    Communications Owner

    An update saying we are still investigating and will update again in an hour is a real update.

  6. 6

    Responder investigates and applies a mitigation to restore service, recording each action taken in the incident channel as it happens.

    Responder

    Restore service first and find the root cause second. Record actions as you go, because reconstructing a timeline afterwards from memory is unreliable and the post-mortem depends on it.

  7. 7

    Incident Lead confirms service is restored and verifies it independently rather than assuming.

    Incident Lead

    Service restored: go to step 8

    Not resolved: go to step 6

  8. 8

    Communications Owner confirms resolution to affected customers.

    Communications Owner

  9. 9

    Incident Lead schedules a post-mortem within five working days, while the detail is still recoverable.

    Incident Lead

    Timing: Within five working days

    Any longer and the timeline gets fuzzy and the urgency evaporates.

  10. 10

    Incident Lead runs a blameless post-mortem producing a timeline, the contributing causes, and what made detection or recovery slower than it should have been.

    Incident Lead

    Blameless in practice, not just in name. The moment a post-mortem assigns fault, people stop volunteering the information that makes it useful.

  11. 11

    Incident Lead records each agreed action with a named owner and a due date.

    Incident Lead

    An action without an owner and a date is not an action. This is where most post-mortems quietly fail.

  12. 12

    Incident Lead updates or creates the procedures the incident revealed to be missing or wrong.

    Incident Lead

    This is the step that stops recurrence. If the incident happened because a recovery procedure did not exist, the post-mortem is not finished until it does.

    Procedures updated and actions assigned: the procedure ends

Change these before you use it

  • Define your own severity levels with concrete criteria in step 3, since 'customer impact' needs to mean something specific in your business.
  • Set the notification window in step 4 to whatever your service commitments actually require.
  • Name your real channels and status page. A procedure that says 'the agreed channel' is not followable at 2am.
  • If you have regulatory breach-notification duties, add them explicitly with their statutory deadlines. This template does not cover them.

This is a starting point, not compliance advice. It is written to be adapted, and a procedure that touches access, money, or customer data needs to match how your business actually operates and whatever rules apply to you. Use it as a first draft to edit, not a policy to adopt.

Make it yours in a couple of minutes

Rather than retyping this and editing it, describe your own version of the process out loud or click through it once, and get a first draft with your actual steps, roles, and systems in it. No account needed to see the result.

No account neededNo credit cardSee the whole SOP before you sign up