End-to-end tests that hold up
End-to-end tests are the only tests that see your system the way a user does. They are also the slowest and most fragile. A small, reliable suite earns trust; a large, flaky one gets ignored.
Choose the journeys
Pick the few flows where a break would hurt the business or users most: sign-up and login, checkout, the core action your product exists for. A good test is a journey someone would notice within minutes if it broke in production.
Leave validation rules, edge cases and permutations to API and unit tests. If you want to check ten discount rules, test them below the UI and keep one end-to-end test proving a discount appears at checkout.
Stable selectors
Tests break most often because they find elements by something that changes: CSS classes, DOM position, generated ids. Prefer, in order:
- roles and accessible names, such as the button named “Place order”;
- labels for form fields;
- test ids (
data-testid) where there is no meaningful role or label.
Role-based selectors survive styling changes and also nudge the markup towards accessibility. If you can’t find a button by its role and name, a screen reader user probably can’t either.
Wait on conditions, not time
A fixed sleep(2000) is either too short on a slow CI machine or wasted time everywhere else. Wait for the state you need: an element visible, a request finished, a URL changed. Playwright’s assertions and Cypress’s commands retry automatically until a timeout, so assert the outcome and let the tool poll.
Set up data through the API
Clicking through five screens to create an order before testing the order page makes the test slow and couples it to screens it isn’t testing. Create preconditions through the API or a seeding endpoint, then use the UI only for the behaviour under test. Do the same for login: sign in once and reuse the stored session.
test("customer can reorder a past order", async ({ page, request }) => {
const order = await createOrder(request, { items: ["SKU-1"] });
await page.goto(`/orders/${order.id}`);
await page.getByRole("button", { name: "Reorder" }).click();
await expect(page.getByRole("heading", { name: "Basket" }))
.toBeVisible();
await expect(page.getByTestId("basket-line")).toHaveCount(1);
});
Isolate every test
Each test creates its own user and data, and never relies on another test having run first. Unique identifiers (an email with a random suffix, for example) let tests run in parallel against the same environment without colliding. Clean up where you can, but design so leftover data can’t break the next run.
Helpers without over-abstraction
Page objects or small helper functions keep selectors in one place, so a changed label is a one-line fix. Keep them thin: methods that describe user actions (checkout.payWithCard()), not deep class hierarchies or a custom language. The test body should still read as a story, with assertions visible in the test rather than hidden inside helpers.
Run in CI with evidence
An end-to-end failure you can’t diagnose is just noise. Configure the runner to keep artefacts from failures:
// playwright.config.ts
export default defineConfig({
retries: process.env.CI ? 1 : 0,
use: {
trace: "on-first-retry",
screenshot: "only-on-failure",
video: "retain-on-failure",
},
});
A Playwright trace records each action with DOM snapshots, network requests and console logs, so you can step through a failure from CI on your own machine. Cypress offers screenshots and videos for the same purpose.
A single retry is useful for capturing a trace, but a test that only passes on retry is flaky. Report it and fix it; don’t let retries hide it.
Habits
- Keep the suite small enough to finish in a few minutes, sharded if needed.
- Select by role, label or test id, never by styling.
- Never use fixed sleeps; assert on the state you expect.
- Create data through the API and give each test its own.
- Save traces, screenshots and logs for every CI failure.
- Delete end-to-end tests that duplicate coverage lower down.