DevOps
Lesson 8 of 8About 3 min readSuggest an edit

Incident response and postmortems

Systems fail no matter how careful you are. What turns a 20-minute blip into a six-hour outage is usually poor coordination: nobody knows who is in charge, what the priority is or who talks to customers. Agree the process in advance, so nobody invents it under pressure.

Severity levels

A severity level tells everyone how serious an incident is and what response it gets. Keep the scale short and define it by impact on users, not by which component broke.

Level Example Response
SEV1 Checkout down for all users, data loss Page immediately, all hands
SEV2 Major feature degraded, many users affected Page on-call, public status update
SEV3 Minor impact, workaround exists Handle in working hours

When unsure, declare the higher severity; downgrading later is cheap.

Roles

For anything beyond a small incident, split the work:

  • The incident commander (IC) coordinates. They decide priorities, assign tasks, and make calls such as “roll back now”. They should not be the person typing commands, because hands-on debugging takes all of your attention.
  • The communications lead writes status updates for customers and internal stakeholders, so the people fixing things are not interrupted with “any news?”.
  • Responders investigate and apply fixes, and report back to the IC.

In a small team one person may hold two roles, but say explicitly who is doing what.

Mitigate first, diagnose later

The first goal is to stop the harm to users, not to understand the cause. Mitigation restores service by the fastest safe means: roll back the last deploy, turn off a feature flag, fail over to another region, scale up, or block an abusive client. Many incidents follow a change, so “what changed recently?” is the first question, and rolling back is often the right first move even before you are sure it is the cause.

Once users are no longer affected, you can find the root cause without time pressure. Keep a timeline in the incident channel as you go: what you saw, what you tried and when.

Status updates

Post updates on a regular cadence, for example every 30 minutes for a major incident, even when there is nothing new. Silence makes customers assume the worst. A useful update is short and factual:

[14:32 UTC] Investigating: some customers see errors at checkout.
We have identified a likely cause and are rolling back a recent
change. Next update by 15:00 UTC.

Don’t guess at causes in public or promise a fix time you can’t back up.

Blameless postmortems

After a significant incident, write a postmortem: what happened, the impact, the timeline, contributing factors and what you will change. Make it blameless. Assume people acted reasonably with the information and tools they had, and ask why the system allowed the mistake. “An engineer ran the wrong command” leads nowhere; “the production and staging commands differ by one flag and nothing asks for confirmation” leads to a fix. People who fear blame hide details.

Incidents rarely have a single root cause, so look for several contributing factors.

Action items that get done

Postmortems are only worth the effort if the follow-up happens. Each action item should be specific, have one owner and a due date, and live in the normal ticket tracker where it is prioritised alongside feature work. If the same kind of incident repeats, check whether earlier actions were completed.

On-call health

On-call only works if it is sustainable. Watch how often people are paged, especially at night, and treat noisy alerts as bugs. Keep rotations large enough for long gaps between shifts, and give time off after a bad night. Exhausted responders make slower, riskier decisions in the next incident.

Checklist

  • Declare the incident, set a severity and name an IC.
  • Mitigate before diagnosing; roll back if a change is suspect.
  • Keep a timeline and post updates on a fixed cadence.
  • Write a blameless postmortem for every significant incident.
  • Give each action item an owner and a date, and track it to completion.

Test what you just read

The DevOps assessment picks questions near your level and explains every answer.