Performance and load testing
Functional tests tell you the system gives the right answer. Performance tests tell you whether it still does so quickly, for many users at once, for hours at a time. These problems usually appear only under load, so you create that load on purpose.
Types of test
| Test | Shape of load | Question it answers |
|---|---|---|
| Load | Expected peak, held steady | Do we meet our goals on a busy day? |
| Stress | Increasing until something breaks | Where is the limit, and how do we fail? |
| Soak | Normal load for many hours | Do memory, connections or disk leak over time? |
| Spike | Sudden jump, then drop | Do autoscaling, queues and caches cope? |
Most teams start with a load test and add the others as specific risks appear.
Goals come from SLOs
“Make it fast” is not a pass criterion. Take goals from your service level objectives, and state them as numbers:
- latency percentiles: 95% of search requests under 300 ms, 99% under 800 ms;
- throughput: sustain 200 requests per second;
- error rate: below 1% under that load.
Write these as thresholds in the test so it passes or fails automatically.
Realistic workloads
A test that hammers one endpoint with identical requests mostly measures your cache. Model real traffic instead:
- a mix of operations in production proportions, for example 70% browse, 20% search, 10% checkout;
- varied data, so requests hit different rows and cache keys;
- think time, the pauses real users make between actions;
- a sensible load model. Many tools default to a fixed number of virtual users that each wait for a response before sending the next request. When the system slows down, they send less, which can hide the problem. For public traffic, a fixed arrival rate is often more realistic.
Example with k6
import http from "k6/http";
import { check, sleep } from "k6";
export const options = {
stages: [
{ duration: "2m", target: 50 }, // ramp up
{ duration: "10m", target: 50 }, // hold
{ duration: "2m", target: 0 }, // ramp down
],
thresholds: {
http_req_duration: ["p(95)<300", "p(99)<800"],
http_req_failed: ["rate<0.01"],
},
};
export default function () {
const q = ["shoes", "coat", "bag"][Math.floor(Math.random() * 3)];
const res = http.get(`${__ENV.BASE_URL}/search?q=${q}`);
check(res, { "status 200": (r) => r.status === 200 });
sleep(1 + Math.random() * 3); // think time
}
If a threshold is breached, k6 exits with a non-zero code, so the run can gate a pipeline.
Environment parity
Results only transfer if the test environment resembles production: similar instance sizes, the same database engine and configuration, realistic data volumes, and the same limits on connections and rate. A query that is instant on 1,000 rows can take seconds on 50 million. Also check the load generator isn’t itself the bottleneck.
Finding bottlenecks
Load tests show symptoms; monitoring shows causes. While the test runs, watch CPU, memory, query times, connection pools, queue depth and downstream calls. Increase load step by step and note where latency starts to climb. The first resource to saturate is your bottleneck. Fix it, re-run, and find the next one.
Read percentiles, not averages
Averages hide the users who suffer. If 95 requests take 100 ms and 5 take 4 seconds, the average is about 300 ms, which looks fine while one in twenty users waits four seconds. Look at p50, p95 and p99, and at the maximum.
Percentiles cannot be averaged across servers or time windows; compute them from the combined raw data or histograms. Compare runs against a baseline, and watch the error rate: a fast error is not a success.
How to decide what to run
- Before a launch or expected traffic peak: a load test at peak, plus a spike test.
- When changing infrastructure or a hot code path: a load test compared against the baseline.
- After leaks or slow degradation in production: a soak test.
- To plan capacity: a stress test to find the breaking point and how the system fails.