Skip to main content
SeleniumDecoded

AI-Assisted Selenium Testing

Where AI genuinely helps a Selenium suite in 2026: generating page objects and tests, self-healing locators, visual assertions, triaging failures, and MCP-driven agents, plus the failure modes to avoid.

Selenium 4 Medium Updated 9 Sept 2026 · Verified against Selenium 4.48.0

Every test tool vendor now advertises AI. Some of it is real and saves hours a week; some of it produces tests that pass while the product is broken. This lesson separates the two, with concrete workflows you can adopt in a Selenium suite today and the guardrails that keep them honest.

Four Places AI Earns Its Keep

Use caseMaturity in 2026Typical tools
Generating page objects and test skeletons from a page or a specHigh; daily useLLM coding assistants (Claude Code, Copilot, Cursor), custom scripts
Self-healing locators when the DOM changesMedium; production use with reviewHealenium, cloud vendor “AI locators”, in-house scoring
Visual and layout assertionsHigh for regression, medium for “does this look right”Applitools, Percy, open-source pixel diff + AI triage
Failure triage and flaky-test clusteringMedium; big time saver on large suitesVendor dashboards, LLM summaries of logs, screenshots and BiDi events

Notably absent: fully autonomous test agents that explore your app and decide what to assert. They exist, they demo well, and in 2026 they still produce tests nobody can maintain. Use them for exploratory coverage reports, not for your regression suite.

1. Generating Page Objects and Tests

The most reliable win. Give an assistant the rendered HTML of a page (or a screenshot plus the DOM) and your project’s page object conventions, and it produces a first draft in seconds. The draft is only as good as the prompt and the DOM, so:

  • Feed it real markup, not a description. driver.getPageSource() or a saved HTML file.
  • State your locator policy: prefer data-testid, then id, then CSS; never absolute XPath.
  • Show one existing page object so it copies your style, wait strategy and naming.
  • Ask for the test data and assertions separately; models are better at structure than at knowing what your business rules are.

A prompt that works:

Here is the DOM of our checkout page (attached) and our BasePage class (attached).
Generate CheckoutPage.java following the same conventions:
- locators as private By fields, data-testid first, then id, then CSS
- public methods named as user actions (enterShippingAddress, selectPaymentMethod, placeOrder)
- every method that changes the page waits with WebDriverWait for the next stable element
- no Thread.sleep, no XPath unless there is no alternative
Return only the Java file.

Then review it like any pull request. In particular check that generated waits wait on the right thing and that assertions are not tautologies.

2. Self-Healing Locators

When a locator stops matching, a self-healing layer finds the “most similar” element using attributes, text, position and DOM structure captured on earlier successful runs, and uses that instead. Healenium is the established open-source option for Java; cloud vendors ship proprietary equivalents.

Healenium wraps the driver (Java); scoring-based fallback pattern (other languages)
Selenium 4 Medium
// Healenium: wrap the driver once, keep the rest of the suite unchanged
import com.epam.healenium.SelfHealingDriver;
WebDriver delegate = new ChromeDriver();
SelfHealingDriver driver = SelfHealingDriver.create(delegate);
// Uses normal locators. On NoSuchElementException, Healenium consults its
// stored DOM snapshots and substitutes the closest match, logging the change.
driver.findElement(By.id("place-order")).click();
// Healenium needs its backend (Docker) to store snapshots:
// docker compose up -d (healenium-backend + postgres)
# No Healenium for Python; a lightweight fallback pattern gives 80% of the value.
from selenium.common.exceptions import NoSuchElementException
def find_with_fallback(driver, primary, *fallbacks):
"""Try the primary locator, then fallbacks in order; log which one worked."""
for by, value in (primary, *fallbacks):
try:
el = driver.find_element(by, value)
if (by, value) != primary:
print(f"HEALED: {primary} -> {(by, value)}")
return el
except NoSuchElementException:
continue
raise NoSuchElementException(f"None of the locators matched: {primary}, {fallbacks}")
place_order = find_with_fallback(
driver,
(By.CSS_SELECTOR, "[data-testid='place-order']"),
(By.ID, "place-order"),
(By.XPATH, "//button[normalize-space()='Place order']"),
)
// Fallback pattern; logs when a secondary locator was needed
async function findWithFallback(driver, primary, ...fallbacks) {
for (const locator of [primary, ...fallbacks]) {
const found = await driver.findElements(locator);
if (found.length) {
if (locator !== primary) console.warn('HEALED:', primary, '->', locator);
return found[0];
}
}
throw new Error(`None of the locators matched: ${primary}`);
}
const placeOrder = await findWithFallback(
driver,
By.css("[data-testid='place-order']"),
By.id('place-order'),
By.xpath("//button[normalize-space()='Place order']")
);
// Fallback pattern; logs when a secondary locator was needed
IWebElement FindWithFallback(IWebDriver driver, By primary, params By[] fallbacks)
{
foreach (var locator in new[] { primary }.Concat(fallbacks))
{
var found = driver.FindElements(locator);
if (found.Count > 0)
{
if (locator != primary) Console.WriteLine($"HEALED: {primary} -> {locator}");
return found[0];
}
}
throw new NoSuchElementException($"None of the locators matched: {primary}");
}
var placeOrder = FindWithFallback(driver,
By.CssSelector("[data-testid='place-order']"),
By.Id("place-order"),
By.XPath("//button[normalize-space()='Place order']"));

The rule that makes self-healing safe: every heal must be logged and reviewed, and the primary locator must be updated in source. A heal that silently persists is a test that no longer checks what you think. Treat “HEALED” log lines as failures in CI on the main branch and as warnings on feature branches.

3. Visual and Layout Assertions

Pixel-diff tools have existed for years; the AI part is ignoring differences that do not matter (anti-aliasing, dynamic dates, ad slots) and flagging the ones that do (overlapping text, missing icons). The Selenium side is unchanged: take a screenshot at a known state and hand it to the service.

Screenshot at a stable state for visual comparison
Selenium 4 Stable
// Wait for the page to settle, then capture the region you care about
wait.until(ExpectedConditions.invisibilityOfElementLocated(By.cssSelector(".skeleton-loader")));
WebElement card = driver.findElement(By.cssSelector("[data-testid='pricing-table']"));
byte[] png = card.getScreenshotAs(OutputType.BYTES);
// Hand to your visual service (Applitools, Percy) or your own baseline comparison
visualService.check("pricing-table", png);
wait.until(EC.invisibility_of_element_located((By.CSS_SELECTOR, ".skeleton-loader")))
card = driver.find_element(By.CSS_SELECTOR, "[data-testid='pricing-table']")
png = card.screenshot_as_png
visual_service.check("pricing-table", png)
await driver.wait(until.elementIsNotVisible(await driver.findElement(By.css('.skeleton-loader'))), 10000);
const card = await driver.findElement(By.css("[data-testid='pricing-table']"));
const png = Buffer.from(await card.takeScreenshot(), 'base64');
await visualService.check('pricing-table', png);
wait.Until(d => !d.FindElements(By.CssSelector(".skeleton-loader")).Any(e => e.Displayed));
var card = driver.FindElement(By.CssSelector("[data-testid='pricing-table']"));
var png = ((ITakesScreenshot)card).GetScreenshot().AsByteArray;
visualService.Check("pricing-table", png);

Mask or freeze dynamic content (dates, counters, carousels) before capturing. BiDi’s emulation commands can pin the viewport and media features so screenshots are deterministic across machines.

4. Failure Triage with Rich Evidence

An LLM cannot tell you why a test failed from a stack trace alone, but it is good at summarising a bundle of evidence: the assertion message, the last screenshot, the page source, the console errors and the network failures. BiDi is what makes that bundle cheap to collect.

Collect evidence per test for AI or human triage
Selenium 4 Stable
// In a JUnit 5 extension or TestNG listener, on failure:
List<String> consoleErrors = new CopyOnWriteArrayList<>();
driver.script().addJavaScriptErrorHandler(e -> consoleErrors.add(e.getText()));
// ... test runs ...
Path dir = Path.of("evidence", testName);
Files.createDirectories(dir);
Files.write(dir.resolve("screenshot.png"), ((TakesScreenshot) driver).getScreenshotAs(OutputType.BYTES));
Files.writeString(dir.resolve("page.html"), driver.getPageSource());
Files.writeString(dir.resolve("console.txt"), String.join("\n", consoleErrors));
Files.writeString(dir.resolve("assertion.txt"), throwable.getMessage());
// Feed the folder to your triage prompt or attach to the CI report
# In a pytest fixture with autouse=True, after yield:
console_errors = []
driver.script.add_javascript_error_handler(lambda e: console_errors.append(e.text))
# ... test runs ...
evidence = Path("evidence") / test_name
evidence.mkdir(parents=True, exist_ok=True)
driver.save_screenshot(evidence / "screenshot.png")
(evidence / "page.html").write_text(driver.page_source)
(evidence / "console.txt").write_text("\n".join(console_errors))
(evidence / "assertion.txt").write_text(str(exc))
const consoleErrors = [];
await driver.script().addJavaScriptErrorHandler((e) => consoleErrors.push(e.text));
// ... test runs ... on failure:
const dir = path.join('evidence', testName);
fs.mkdirSync(dir, { recursive: true });
fs.writeFileSync(path.join(dir, 'screenshot.png'), await driver.takeScreenshot(), 'base64');
fs.writeFileSync(path.join(dir, 'page.html'), await driver.getPageSource());
fs.writeFileSync(path.join(dir, 'console.txt'), consoleErrors.join('\n'));
fs.writeFileSync(path.join(dir, 'assertion.txt'), err.message);
var consoleErrors = new ConcurrentBag<string>();
var bidi = await driver.AsBiDiAsync();
await bidi.Log.OnEntryAddedAsync(e => { if (e.Level == "error") consoleErrors.Add(e.Text); });
// ... test runs ... on failure:
var dir = Path.Combine("evidence", testName);
Directory.CreateDirectory(dir);
((ITakesScreenshot)driver).GetScreenshot().SaveAsFile(Path.Combine(dir, "screenshot.png"));
File.WriteAllText(Path.Combine(dir, "page.html"), driver.PageSource);
File.WriteAllText(Path.Combine(dir, "console.txt"), string.Join("\n", consoleErrors));
File.WriteAllText(Path.Combine(dir, "assertion.txt"), ex.Message);

A triage prompt then asks: “Given this assertion, screenshot, console errors and network failures, is this most likely a product bug, a test timing issue, or an environment problem? Cite the evidence.” Cluster the answers across a nightly run and flaky tests separate themselves from real regressions.

Agents and MCP

Model Context Protocol (MCP) servers let an AI assistant drive a real browser through Selenium or a Grid, read the DOM, and run your test suite. In 2026 the practical uses are:

  • Reproducing a bug report: “open the staging site, log in as the test user, follow these steps, screenshot each step.”
  • Drafting a test from an exploratory session: the agent records the actions it took, then writes a page-object-based test you review.
  • Answering “what does this page look like now?” during code review.

Keep agents off production data, give them a dedicated test account, and never let them commit generated tests without a human reviewing the assertions.

Failure Modes to Watch For

  • Tautological assertions: generated tests that assert an element exists after clicking it. Ask “what would this fail on?” for every assertion.
  • Silent healing: covered above. Log, review, fix the source.
  • Prompt drift: page objects generated with different prompts diverge in style. Keep the prompt in the repository next to the conventions doc.
  • Over-mocking: AI-generated network mocks that describe the API as the model imagines it. Record real responses with BiDi first.
  • Confidence without evidence: an LLM summary of a failure is a hypothesis. The screenshot and console log are the evidence.

Summary

  • Use AI for drafting page objects and tests, reviewing them as you would a junior engineer’s pull request.
  • Self-healing locators are a maintenance aid, not a substitute for fixing locators; log and review every heal.
  • Visual AI removes noise from screenshot comparison; you still choose stable states and mask dynamic content.
  • BiDi makes evidence collection cheap, which makes AI triage useful.
  • Autonomous test generation is a demo, not a regression strategy, in 2026.

Copy-paste recipes for this topic

Related lessons