Flaky Tests: Diagnosis and Cure
A systematic method for finding why a test passes sometimes: classify the failure, reproduce it deliberately, fix the root cause (timing, state, order, environment, product), and quarantine without hiding.
A flaky test is one that passes and fails without any change to the code under test. One flaky test is an annoyance; ten make the team stop trusting the suite, and an untrusted suite is worse than none because real failures are dismissed as noise. Flakiness is not random. Every flaky test has a deterministic cause that only looks random because the trigger is timing or state you do not control. This lesson gives you the method to find it.
Step 1: Classify the Failure
Gather the evidence from the last five failures (screenshots, stack traces, console logs, timings) and put each into one of five buckets:
| Bucket | Signature | Typical root cause |
|---|---|---|
| Timing | NoSuchElement, StaleElementReference, ElementClickIntercepted, assertion on old text | Missing or wrong wait; app renders asynchronously |
| State leakage | Fails only after another specific test; different data than expected | Shared user, shared database rows, cookies, browser reuse |
| Order or parallelism | Passes alone, fails in the full run | Shared static driver, shared fixtures, port collisions |
| Environment | Fails only in CI or only on one machine | Window size, fonts, network speed, browser version, /dev/shm |
| Product | Intermittent real bug | Race condition in the app, backend timeout |
The fifth bucket is the one everyone forgets. Around a fifth of “flaky” failures in mature suites are intermittent product defects the test is correctly catching.
Step 2: Reproduce Deliberately
You cannot fix what you cannot reproduce. Make the failure likely instead of waiting for it.
// JUnit 5: repeat the suspect test@RepeatedTest(50)void appliesCoupon() { /* ... */ }
// Add artificial latency to reveal timing bugs (BiDi, Chromium and Firefox)// or via CDP on Chromium:driver.execute_cdp_cmd — see Python; in Java use Network.emulateNetworkConditions from DevTools
// Slow the CPU so renders take longer (CDP)devTools.send(Emulation.setCPUThrottlingRate(4));# pytest-repeat: pytest tests/test_coupon.py --count=50 -x# pytest-randomly: shuffle test order to expose state leakage# pytest -p randomly --randomly-seed=12345
# Make timing bugs likely: throttle the network and CPUdriver.execute_cdp_cmd("Network.enable", {})driver.execute_cdp_cmd("Network.emulateNetworkConditions", {"offline": False, "latency": 300, "downloadThroughput": 50_000, "uploadThroughput": 20_000})driver.execute_cdp_cmd("Emulation.setCPUThrottlingRate", {"rate": 4})// mocha: --retries is NOT the tool; use a loop in a scratch filefor (let i = 0; i < 50; i++) {it(`applies coupon (run ${i})`, async () => { /* ... */ });}
// Throttle to reveal timing bugs (Chromium)await driver.sendDevToolsCommand('Network.enable', {});await driver.sendDevToolsCommand('Network.emulateNetworkConditions',{ offline: false, latency: 300, downloadThroughput: 50000, uploadThroughput: 20000 });await driver.sendDevToolsCommand('Emulation.setCPUThrottlingRate', { rate: 4 });// NUnit: [Repeat(50)] on the suspect test[Test, Repeat(50)]public void AppliesCoupon() { /* ... */ }
// Throttle via CDP (Chromium)var session = ((ChromeDriver)driver).GetDevToolsSession();await session.SendCommand("Network.enable", new JsonObject());await session.SendCommand("Emulation.setCPUThrottlingRate", JsonSerializer.SerializeToNode(new { rate = 4 }));If it fails under throttling, it is a timing bug and you have a reliable reproduction. If it only fails in the full suite, shuffle test order; if the failure follows a specific predecessor, it is state leakage. If it never fails locally, diff the environments.
Step 3: Fix the Root Cause
Timing
The fix is always a wait on the right condition. Not a longer implicit wait, not a sleep.
// FLAKY: waits for presence, but the element is re-rendered after data loadsWebElement total = wait.until(ExpectedConditions.presenceOfElementLocated(By.id("total")));driver.findElement(By.id("apply-coupon")).click();assertEquals("$89.00", total.getText()); // stale, or reads the pre-coupon value
// ROBUST: wait for the state you actually assert ondriver.findElement(By.id("apply-coupon")).click();wait.until(ExpectedConditions.textToBe(By.id("total"), "$89.00"));
// ROBUST alternative when the value is unknown: wait for it to change, then stabiliseString before = driver.findElement(By.id("total")).getText();driver.findElement(By.id("apply-coupon")).click();wait.until(d -> !d.findElement(By.id("total")).getText().equals(before));# FLAKYtotal = wait.until(EC.presence_of_element_located((By.ID, "total")))driver.find_element(By.ID, "apply-coupon").click()assert total.text == "$89.00"
# ROBUSTdriver.find_element(By.ID, "apply-coupon").click()wait.until(EC.text_to_be_present_in_element((By.ID, "total"), "$89.00"))
# ROBUST when the value is unknownbefore = driver.find_element(By.ID, "total").textdriver.find_element(By.ID, "apply-coupon").click()wait.until(lambda d: d.find_element(By.ID, "total").text != before)// FLAKYconst total = await driver.wait(until.elementLocated(By.id('total')), 10000);await driver.findElement(By.id('apply-coupon')).click();assert.strictEqual(await total.getText(), '$89.00');
// ROBUSTawait driver.findElement(By.id('apply-coupon')).click();await driver.wait(until.elementTextIs(driver.findElement(By.id('total')), '$89.00'), 10000);// FLAKYvar total = wait.Until(d => d.FindElement(By.Id("total")));driver.FindElement(By.Id("apply-coupon")).Click();Assert.That(total.Text, Is.EqualTo("$89.00"));
// ROBUSTdriver.FindElement(By.Id("apply-coupon")).Click();wait.Until(d => d.FindElement(By.Id("total")).Text == "$89.00");Common timing culprits and their conditions:
- Clicked too early:
elementToBeClickable, plus wait for overlays to disappear. - Read too early:
textToBe,attributeToBe, or a custom “value changed” condition. - Navigated too early:
urlContainsafter the click. - Animations: disable them with a preload script (see WebDriver BiDi) or wait for
getComputedStyleto settle. - API calls: a BiDi network-idle wait (see Custom Wait Conditions).
State leakage
- Give every test its own user or tenant, created via API in setup, deleted in teardown.
- Fresh driver per test. Reusing a browser to save two seconds costs hours of debugging.
- Never depend on test order. Randomise it in CI so dependencies fail loudly.
- Clear cookies and storage if you must reuse a session.
Parallelism
- No static driver; use a ThreadLocal or per-test fixture. See Driver Factory and Thread Safety.
- Unique file names for downloads and screenshots per test.
- Backend rate limits and connection pools sized for the parallelism.
Environment
- Explicit window size in every driver factory.
- Pin browser versions with Selenium Manager in CI.
--shm-size=2gfor Docker.- Install fonts in Linux images if layout matters.
Product
Log it as a bug with the evidence, mark the test with the bug id, and keep it running. Do not “fix” the test by waiting longer; you would be hiding a defect.
Step 4: Quarantine Without Hiding
While a fix is in progress, a flaky test should neither block the pipeline nor disappear. Tag it, run it in a separate non-blocking job, and track it.
// JUnit 5@Tag("flaky")@Testvoid appliesCouponUnderLoad() { /* ... */ }
// Main job: mvn test -Dgroups='!flaky'// Quarantine job: mvn test -Dgroups=flaky (allowed to fail, results tracked)@pytest.mark.flaky(reason="JIRA-1234 race in coupon service")def test_applies_coupon_under_load(driver): ...
# Main job: pytest -m "not flaky"# Quarantine job: pytest -m flaky --continue-on-collection-errors (allowed to fail)# Register the marker in pytest.ini: markers = flaky: known flaky, tracked in JIRA// mocha grep on a title conventionit('applies coupon under load @flaky', async () => { /* ... */ });
// Main job: mocha --grep @flaky --invert// Quarantine job: mocha --grep @flaky[Test, Category("Flaky")]public void AppliesCouponUnderLoad() { /* ... */ }
// Main job: dotnet test --filter "TestCategory!=Flaky"// Quarantine job: dotnet test --filter "TestCategory=Flaky"Set a rule: a test may stay in quarantine for two weeks. After that it is fixed or deleted.
The Retry Question
Automatic retries (@RetryingTest, pytest-rerunfailures, --retries) make a green pipeline out of a flaky suite and are the fastest way to lose visibility. If you use them at all:
- Only on tests tagged flaky, never globally.
- Report retries as a metric; a test that needs retries is failing.
- Never retry a test that has side effects on shared state; the second run starts from a different state than the first.
Measuring Flakiness
Track per test: total runs, failures, and failures that passed on retry. A dashboard (Allure history, your CI’s test analytics, or a spreadsheet from JUnit XML) makes the top ten flaky tests obvious. Fix those first; they usually share a cause.
A Debugging Session, Start to Finish
test_checkout_applies_couponfailed 3 times in 40 nightly runs withAssertionError: expected $89.00 got $99.00.- Screenshot shows the coupon banner “Applied!” but the old total. Timing bucket.
- Repeat 50 times locally with CPU throttling: 12 failures. Reproduced.
- The test waited for the banner, then read the total. The total updates from a second API call.
- Fix: wait for
textToBe("#total", "$89.00"). 200 throttled runs, 0 failures. - Remove the flaky tag, close the ticket.
Summary
- Flakiness has deterministic causes: timing, state, order, environment, or a real product bug.
- Reproduce with repetition, throttling and shuffling before touching the code.
- Fix the cause with the right wait, isolated state, thread-safe drivers or pinned environments.
- Quarantine with tags and a deadline; retries hide problems rather than solve them.