Testing in production safely
Pre-production testing reduces risk; it cannot remove it. Some bugs only appear with real users, real data and real traffic. Testing in production doesn’t replace earlier testing. It adds controlled ways to find those bugs early, while the impact is still small.
Why staging never matches production
Staging differs in ways that matter:
- data: smaller, cleaner and missing the odd records years of real use create;
- traffic: lower volume and none of the strange clients, retries and bots;
- configuration: different secrets, limits, feature settings and third-party accounts;
- scale: fewer instances, so no contention, cache eviction or cross-zone latency.
So the question is not whether production will surprise you, but how quickly you notice and how few users are affected.
Feature flags and dark launches
A feature flag separates deploying code from releasing it. You ship the code switched off, then enable it for internal staff, then a small share of users, then everyone. Turning the flag off is faster and safer than a rollback.
A dark launch runs new code on real traffic without showing the result. For example, a new pricing service receives a copy of each request, and its answer is compared with the old service’s and logged. You learn about correctness and load before any user depends on it. Be careful that the shadow path has no side effects: no emails, no charges, no writes to shared data.
Canary releases
A canary deploys the new version to a small slice of servers or traffic, then compares its health with the current version: error rate, latency percentiles, key business events. If the canary is worse, traffic moves back automatically. If it is healthy, the rollout widens in steps. Compare against the baseline running at the same time, not against yesterday, so normal daily changes in traffic don’t mislead you.
Synthetic monitoring
Synthetic monitoring runs scripted journeys against production on a schedule, often from several regions. It catches breakage at 3 a.m. when no real user is there to hit it, and it gives a steady signal that doesn’t depend on traffic.
// Runs every 5 minutes against production
test("synthetic: search and add to basket", async ({ page }) => {
await loginAs(page, process.env.SYNTHETIC_USER!);
await page.goto("/search?q=umbrella");
await page.getByRole("link", { name: /umbrella/i }).first().click();
await page.getByRole("button", { name: "Add to basket" }).click();
await expect(page.getByTestId("basket-count")).toHaveText("1");
await clearBasket(page);
});
Keep synthetic checks few and stable, and alert on repeated failures, not a single blip.
Observability as a testing tool
In production, your assertions are your telemetry. Structured logs, metrics and traces let you ask “did the new code path error, and for whom?” Before enabling a flag, decide which signals would show it working or failing, and make sure they exist. A release you can’t observe is a release you can’t test.
A/B tests are not correctness tests
An A/B test measures whether a change improves an outcome, such as conversion. It assumes both variants work. A broken variant can look like a losing variant, so check correctness first with flags, canaries and monitoring, then run the experiment.
Guardrails
Testing in production is only responsible with safety measures in place:
- kill switches: every risky feature can be turned off in seconds without a deploy;
- rate limits and blast radius: start with a small percentage, one region or internal users;
- test accounts: clearly marked, excluded from analytics, billing and customer emails;
- data cleanup: tag test data so it can be found and removed, and never leave it in reports;
- automatic rollback: tie canary and flag rollouts to error and latency thresholds;
- privacy: shadow traffic and logs follow the same data protection rules as the real thing.
Checklist before a production rollout
- Is the change behind a flag, with a named owner who can switch it off?
- Are success and failure signals defined, with dashboards and alerts?
- Does the rollout start small and widen only when metrics are healthy?
- Are synthetic checks covering the affected journey?
- Is all test data and traffic marked, isolated and cleaned up?