QA
Lesson 6 of 8About 3 min readSuggest an edit

Performance and load testing

Functional tests tell you the system gives the right answer. Performance tests tell you whether it still does so quickly, for many users at once, for hours at a time. These problems usually appear only under load, so you create that load on purpose.

Types of test

Test Shape of load Question it answers
Load Expected peak, held steady Do we meet our goals on a busy day?
Stress Increasing until something breaks Where is the limit, and how do we fail?
Soak Normal load for many hours Do memory, connections or disk leak over time?
Spike Sudden jump, then drop Do autoscaling, queues and caches cope?

Most teams start with a load test and add the others as specific risks appear.

Goals come from SLOs

“Make it fast” is not a pass criterion. Take goals from your service level objectives, and state them as numbers:

  • latency percentiles: 95% of search requests under 300 ms, 99% under 800 ms;
  • throughput: sustain 200 requests per second;
  • error rate: below 1% under that load.

Write these as thresholds in the test so it passes or fails automatically.

Realistic workloads

A test that hammers one endpoint with identical requests mostly measures your cache. Model real traffic instead:

  • a mix of operations in production proportions, for example 70% browse, 20% search, 10% checkout;
  • varied data, so requests hit different rows and cache keys;
  • think time, the pauses real users make between actions;
  • a sensible load model. Many tools default to a fixed number of virtual users that each wait for a response before sending the next request. When the system slows down, they send less, which can hide the problem. For public traffic, a fixed arrival rate is often more realistic.

Example with k6

import http from "k6/http";
import { check, sleep } from "k6";

export const options = {
  stages: [
    { duration: "2m", target: 50 },  // ramp up
    { duration: "10m", target: 50 }, // hold
    { duration: "2m", target: 0 },   // ramp down
  ],
  thresholds: {
    http_req_duration: ["p(95)<300", "p(99)<800"],
    http_req_failed: ["rate<0.01"],
  },
};

export default function () {
  const q = ["shoes", "coat", "bag"][Math.floor(Math.random() * 3)];
  const res = http.get(`${__ENV.BASE_URL}/search?q=${q}`);
  check(res, { "status 200": (r) => r.status === 200 });
  sleep(1 + Math.random() * 3); // think time
}

If a threshold is breached, k6 exits with a non-zero code, so the run can gate a pipeline.

Environment parity

Results only transfer if the test environment resembles production: similar instance sizes, the same database engine and configuration, realistic data volumes, and the same limits on connections and rate. A query that is instant on 1,000 rows can take seconds on 50 million. Also check the load generator isn’t itself the bottleneck.

Finding bottlenecks

Load tests show symptoms; monitoring shows causes. While the test runs, watch CPU, memory, query times, connection pools, queue depth and downstream calls. Increase load step by step and note where latency starts to climb. The first resource to saturate is your bottleneck. Fix it, re-run, and find the next one.

Read percentiles, not averages

Averages hide the users who suffer. If 95 requests take 100 ms and 5 take 4 seconds, the average is about 300 ms, which looks fine while one in twenty users waits four seconds. Look at p50, p95 and p99, and at the maximum.

Percentiles cannot be averaged across servers or time windows; compute them from the combined raw data or histograms. Compare runs against a baseline, and watch the error rate: a fast error is not a success.

How to decide what to run

  • Before a launch or expected traffic peak: a load test at peak, plus a spike test.
  • When changing infrastructure or a hot code path: a load test compared against the baseline.
  • After leaks or slow degradation in production: a soak test.
  • To plan capacity: a stress test to find the breaking point and how the system fails.

Next: Testing in production safely

Feature flags, dark launches, canaries, synthetic monitoring and the guardrails that make production testing safe.