Incident response and postmortems
Systems fail no matter how careful you are. What turns a 20-minute blip into a six-hour outage is usually poor coordination: nobody knows who is in charge, what the priority is or who talks to customers. Agree the process in advance, so nobody invents it under pressure.
Severity levels
A severity level tells everyone how serious an incident is and what response it gets. Keep the scale short and define it by impact on users, not by which component broke.
| Level | Example | Response |
|---|---|---|
| SEV1 | Checkout down for all users, data loss | Page immediately, all hands |
| SEV2 | Major feature degraded, many users affected | Page on-call, public status update |
| SEV3 | Minor impact, workaround exists | Handle in working hours |
When unsure, declare the higher severity; downgrading later is cheap.
Roles
For anything beyond a small incident, split the work:
- The incident commander (IC) coordinates. They decide priorities, assign tasks, and make calls such as “roll back now”. They should not be the person typing commands, because hands-on debugging takes all of your attention.
- The communications lead writes status updates for customers and internal stakeholders, so the people fixing things are not interrupted with “any news?”.
- Responders investigate and apply fixes, and report back to the IC.
In a small team one person may hold two roles, but say explicitly who is doing what.
Mitigate first, diagnose later
The first goal is to stop the harm to users, not to understand the cause. Mitigation restores service by the fastest safe means: roll back the last deploy, turn off a feature flag, fail over to another region, scale up, or block an abusive client. Many incidents follow a change, so “what changed recently?” is the first question, and rolling back is often the right first move even before you are sure it is the cause.
Once users are no longer affected, you can find the root cause without time pressure. Keep a timeline in the incident channel as you go: what you saw, what you tried and when.
Status updates
Post updates on a regular cadence, for example every 30 minutes for a major incident, even when there is nothing new. Silence makes customers assume the worst. A useful update is short and factual:
[14:32 UTC] Investigating: some customers see errors at checkout.
We have identified a likely cause and are rolling back a recent
change. Next update by 15:00 UTC.
Don’t guess at causes in public or promise a fix time you can’t back up.
Blameless postmortems
After a significant incident, write a postmortem: what happened, the impact, the timeline, contributing factors and what you will change. Make it blameless. Assume people acted reasonably with the information and tools they had, and ask why the system allowed the mistake. “An engineer ran the wrong command” leads nowhere; “the production and staging commands differ by one flag and nothing asks for confirmation” leads to a fix. People who fear blame hide details.
Incidents rarely have a single root cause, so look for several contributing factors.
Action items that get done
Postmortems are only worth the effort if the follow-up happens. Each action item should be specific, have one owner and a due date, and live in the normal ticket tracker where it is prioritised alongside feature work. If the same kind of incident repeats, check whether earlier actions were completed.
On-call health
On-call only works if it is sustainable. Watch how often people are paged, especially at night, and treat noisy alerts as bugs. Keep rotations large enough for long gaps between shifts, and give time off after a bad night. Exhausted responders make slower, riskier decisions in the next incident.
Checklist
- Declare the incident, set a severity and name an IC.
- Mitigate before diagnosing; roll back if a change is suspect.
- Keep a timeline and post updates on a fixed cadence.
- Write a blameless postmortem for every significant incident.
- Give each action item an owner and a date, and track it to completion.