Security
Lesson 8 of 8About 4 min readSuggest an edit

Responding to security incidents

Every organisation that runs software will eventually have a security incident: a leaked key, a compromised account, an exploited dependency. The difference between a bad day and a disaster is mostly decided before it happens, by what you logged, who you can reach and whether anyone has practised.

Preparation

Prepare while things are calm:

  • Logging: record logins, permission changes, admin actions, sensitive data access and deploys, with timestamps and user ids. Ship logs where an attacker on one host cannot delete them, and keep them for months, not days.
  • Contacts: a current list of who decides, who owns each system, and how to reach your cloud provider, legal adviser and key vendors out of hours.
  • Runbooks: short, tested steps for likely scenarios such as a leaked cloud key or a compromised employee account. Include how to rotate each type of credential.
  • Roles: name an incident lead who coordinates and decides, separate from the people doing hands-on investigation.

Rehearse with a tabletop exercise: walk through a scenario and note every question nobody could answer.

Detection signals

Typical signals include:

  • secret-scanning alerts or a provider notifying you that a key appeared publicly;
  • logins from new locations followed by permission changes;
  • unusual volumes of data export or database reads;
  • new cloud resources, users or access keys nobody created.

Treat reports from customers and researchers as real until shown otherwise; they are often your earliest warning.

Containment, eradication and recovery

These stages overlap, but each has a different goal.

Stage Goal Examples
Containment stop the damage spreading disable accounts, isolate hosts, revoke tokens
Eradication remove the attacker’s access patch the flaw, rebuild compromised machines
Recovery return to normal service safely restore clean backups, monitor for recurrence

Contain quickly, even imperfectly. Eradicate thoroughly: if you fix the entry point but miss a second access key the attacker created, they are still in.

Preserving evidence

Evidence tells you what happened and what data was affected, and may be needed for legal or insurance purposes. Before destroying anything:

  • snapshot disks and memory of affected machines, rather than terminating them;
  • export relevant logs to a separate, write-protected location;
  • keep a timeline of what you observed and did, with timestamps and names.
14:02 Alert: access key AKIA...7Q used from unknown ASN
14:09 Key disabled (containment), incident lead: Priya
14:40 Found second key created 13:51 by same identity

Rotating credentials

Assume any secret that was reachable from a compromised system is known to the attacker. That includes environment variables, config files and CI secrets. Rotate them in a deliberate order: revoke the credentials that let the attacker create new ones first, then the rest, then invalidate sessions so stolen cookies stop working. Confirm that old credentials actually fail.

Communicating

Keep internal updates regular and factual, in one channel. Externally, tell affected users what happened, what data was involved, what you have done and what they should do, such as resetting a password.

Laws and contracts may require notifying regulators or customers of personal data breaches within strict time limits, sometimes 72 hours. Involve legal and privacy advisers early so deadlines are not missed.

Vulnerability disclosure

Publish a vulnerability disclosure policy that states how to report, what is in scope, how quickly you will respond, and that good-faith research will not be treated as an attack. A file at /.well-known/security.txt tells people where to send reports. Acknowledge reports promptly, fix them, and agree a disclosure date with the reporter.

Learning afterwards

Hold a blameless post-incident review within a week or two. Establish the timeline, the causes, and what helped or slowed you down. Turn the findings into owned, tracked actions: a missing log, a runbook step, a detection rule, a design change.

Checklist

  • Logs for auth, admin and data access, stored out of attackers’ reach.
  • An out-of-hours contact list and named incident lead.
  • Runbooks for rotating every credential type, tested.
  • Snapshots and a written timeline before cleanup.
  • A disclosure policy and security.txt published.

Test what you just read

The Security assessment picks questions near your level and explains every answer.