Threat modelling
Threat modelling is a structured way of asking what can go wrong with a design before you build it. A missing check found on a whiteboard costs a conversation; the same flaw found in production costs an incident. The engineers who design a feature know its data flows best, so they should take part.
A widely used framing is four questions:
- What are we working on?
- What can go wrong?
- What are we going to do about it?
- Did we do a good enough job?
Draw the system
Start with a data-flow diagram: the processes, data stores, external parties and the flows between them. One page of boxes and arrows is enough.
Then mark the trust boundaries: places where data or control passes between parts with different levels of trust. Typical boundaries sit between the browser and your API, your service and a third-party webhook, your application and the database, and one tenant’s data and another’s. Most threats cluster at these lines, because that is where input must be validated and identity checked.
[Browser] --HTTPS--> | [API] ---> [Orders DB]
| |
trust boundary | +---> [Payment provider]
| (external, trusted by contract)
[Webhook sender] --> | [Webhook handler]
Find threats with STRIDE
STRIDE, developed at Microsoft, is a checklist to walk each element and flow through. Each letter is the violation of a security property.
| Threat | Property broken | Example |
|---|---|---|
| Spoofing | authenticity | forged webhook calls with no signature check |
| Tampering | integrity | client edits the price field in a checkout request |
| Repudiation | non-repudiation | admin deletes records and no audit log exists |
| Information disclosure | confidentiality | error pages print stack traces and config |
| Denial of service | availability | unbounded upload size fills the disk |
| Elevation of privilege | authorisation | a user sets role=admin via mass assignment |
Not every letter applies to every element. Aim for breadth first; a list of plausible threats beats a deep analysis of one.
Rate and prioritise
You will find more threats than you can fix. Rate each one on two rough scales:
- likelihood: how reachable is it, and how much skill or access does it need?
- impact: what data, money or availability is lost, and for how many users?
A simple high, medium, low grid on both is usually enough to sort the list. Resist precise scores; the value is in the ranking and the discussion, not the arithmetic. Each threat then gets one of four responses: mitigate it, eliminate it by removing the feature or data, transfer it (to a provider or insurer), or accept it with a named owner and a reason.
Turn threats into requirements and tests
A threat model that ends as a document nobody reads has failed. Convert each mitigated threat into something that lives in the normal workflow:
- a requirement on the ticket: “webhook handler rejects requests without a valid signature”;
- a test that proves it: send an unsigned request and expect 401;
- a detection where prevention is partial: alert on repeated signature failures.
Passing tests are also your answer to the fourth question.
Keep it lightweight
A threat model does not need a week-long workshop. For most features, 30 to 60 minutes with the people building it is enough:
- run it during design, when changes are cheap;
- update it when the design changes a trust boundary, adds a data store or introduces a new external party;
- keep the diagram and threat list next to the code or design doc, not in a separate archive.
How to decide when to do one
Do a threat model when a change:
- handles authentication, payments, personal data or secrets;
- adds a new external integration or public endpoint;
- moves data across a trust boundary for the first time;
- changes who can access what.
For small changes inside an existing boundary, a quick STRIDE pass during code review is usually enough.