Skip to main content
SeleniumDecoded

CI/CD Pipelines: GitHub Actions, GitLab CI and Jenkins

Run Selenium suites in CI reliably: headless browsers, Selenium Manager caching, Grid services in pipelines, parallel jobs, artifacts on failure, and the settings that stop CI-only flakiness.

Selenium 4 Stable Updated 9 Sept 2026 · Verified against Selenium 4.48.0

A Selenium suite that only runs on a laptop is a demo. CI is where it earns its keep, and CI is also where most suites first turn flaky: different screen sizes, no GPU, slower disks, cold caches, and a browser version that changed overnight. This lesson gives you working pipeline definitions for the three most common CI systems and the settings that make them stable.

The Checklist for Any CI System

  1. Headless browser with an explicit window size.
  2. Pinned or cached browser and driver so runs are reproducible and fast. Selenium Manager handles the matching; you handle the cache.
  3. Parallelism matched to the runner’s CPU and memory (roughly one browser per 2 vCPUs and 2 GB).
  4. Artifacts on failure: screenshots, page source, console logs, videos if using Grid.
  5. Timeouts at every level: test, job, and driver, so a hung browser fails fast.
  6. Retry policy for known-flaky tags only, never blanket retries.
  7. Secrets for test accounts and cloud credentials from the CI secret store, never in the repository.

GitHub Actions

Ubuntu runners ship with Chrome, Firefox and Edge preinstalled, and Selenium Manager finds them. Cache its download directory so drivers are not fetched every run.

.github/workflows/e2e.yml
name: E2E
on:
pull_request:
schedule:
- cron: '0 3 * * *' # nightly full run
jobs:
selenium:
runs-on: ubuntu-latest
timeout-minutes: 30
strategy:
fail-fast: false
matrix:
browser: [chrome, firefox]
shard: [1, 2, 3]
env:
BROWSER: ${{ matrix.browser }}
HEADLESS: 'true'
BASE_URL: https://staging.example.com
SE_CACHE_PATH: ${{ github.workspace }}/.se-cache
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with: { distribution: temurin, java-version: '21', cache: maven }
- name: Cache Selenium Manager downloads
uses: actions/cache@v4
with:
path: .se-cache
key: se-${{ runner.os }}-${{ matrix.browser }}-${{ hashFiles('pom.xml') }}
- name: Run shard ${{ matrix.shard }}
run: mvn -B test -Dshard=${{ matrix.shard }} -Dshards=3
env:
TEST_USER: ${{ secrets.TEST_USER }}
TEST_PASSWORD: ${{ secrets.TEST_PASSWORD }}
- name: Upload failure evidence
if: failure()
uses: actions/upload-artifact@v4
with:
name: evidence-${{ matrix.browser }}-${{ matrix.shard }}
path: |
target/screenshots/
target/surefire-reports/
retention-days: 7

The same shape works for Python (actions/setup-python and pytest -n 2 --dist loadfile), Node (actions/setup-node and npx mocha --parallel) and .NET (actions/setup-dotnet and dotnet test).

Running Against a Grid Service Container

When you need a specific browser version or want video, run the official Selenium image as a service and point GRID_URL at it.

services:
selenium:
image: selenium/standalone-chrome:4.48.0
ports: ['4444:4444']
options: >-
--shm-size=2g
--health-cmd "curl -sf http://localhost:4444/status || exit 1"
--health-interval 5s --health-timeout 5s --health-retries 20
env:
GRID_URL: http://localhost:4444

--shm-size=2g matters: Chrome crashes with “session deleted because of page crash” when /dev/shm is the default 64 MB.

GitLab CI

GitLab’s services keyword works the same way. The job container needs your language runtime; the browser lives in the service.

.gitlab-ci.yml
e2e:
stage: test
image: python:3.12
services:
- name: selenium/standalone-chrome:4.48.0
alias: selenium
variables:
SE_NODE_MAX_SESSIONS: "2"
SE_VNC_NO_PASSWORD: "1"
variables:
GRID_URL: http://selenium:4444
BASE_URL: https://staging.example.com
HEADLESS: "true"
FF_NETWORK_PER_BUILD: "true" # services and job share a network
cache:
key: pip-$CI_COMMIT_REF_SLUG
paths: [.cache/pip]
before_script:
- pip install -r requirements.txt
- |
for i in $(seq 1 30); do
curl -sf http://selenium:4444/status | grep -q '"ready": true' && break
sleep 2
done
script:
- pytest tests/e2e -n 2 --junitxml=report.xml
artifacts:
when: always
reports:
junit: report.xml
paths: [failures/]
expire_in: 1 week
parallel: 3 # GitLab sets CI_NODE_INDEX / CI_NODE_TOTAL for sharding
timeout: 30 minutes

Sharding with parallel: 3 requires your test runner to split by CI_NODE_INDEX; pytest can do it with the pytest-split plugin, or partition by file name hash in a conftest.py hook.

Jenkins

A declarative pipeline running tests against a Docker-hosted Grid, with the Grid started in a sidecar and torn down in post.

// Jenkinsfile
pipeline {
agent { label 'docker' }
options { timeout(time: 40, unit: 'MINUTES'); timestamps() }
environment {
GRID_URL = 'http://localhost:4444'
BASE_URL = 'https://staging.example.com'
HEADLESS = 'true'
TEST_CREDS = credentials('e2e-test-user') // exposes TEST_CREDS_USR and TEST_CREDS_PSW
}
stages {
stage('Grid up') {
steps {
sh '''
docker run -d --name grid --shm-size=2g -p 4444:4444 \
-e SE_NODE_MAX_SESSIONS=4 selenium/standalone-chrome:4.48.0
for i in $(seq 1 30); do
curl -sf http://localhost:4444/status | grep -q '"ready": true' && exit 0
sleep 2
done
exit 1
'''
}
}
stage('Test') {
steps {
sh 'mvn -B test -Dparallel.threads=4'
}
}
}
post {
always {
junit 'target/surefire-reports/*.xml'
archiveArtifacts artifacts: 'target/screenshots/**', allowEmptyArchive: true
sh 'docker logs grid > grid.log 2>&1 || true; docker rm -f grid || true'
archiveArtifacts artifacts: 'grid.log', allowEmptyArchive: true
}
}
}

Sharding Tests Across Jobs

Parallel jobs need a deterministic way to split tests. A stable approach in any language: sort test identifiers, then take every Nth.

Select this job's shard of tests
Selenium 4 Stable
// JUnit 5: a condition that skips tests not in this shard
public class ShardCondition implements ExecutionCondition {
@Override
public ConditionEvaluationResult evaluateExecutionCondition(ExtensionContext ctx) {
int shard = Integer.getInteger("shard", 1);
int shards = Integer.getInteger("shards", 1);
String id = ctx.getUniqueId();
int bucket = Math.floorMod(id.hashCode(), shards) + 1;
return bucket == shard
? ConditionEvaluationResult.enabled("in shard " + shard)
: ConditionEvaluationResult.disabled("not in shard " + shard);
}
}
// Register globally via META-INF/services/org.junit.jupiter.api.extension.Extension
// with -Djunit.jupiter.extensions.autodetection.enabled=true
conftest.py
import os
def pytest_collection_modifyitems(config, items):
shard = int(os.getenv("CI_NODE_INDEX", "1"))
shards = int(os.getenv("CI_NODE_TOTAL", "1"))
if shards == 1:
return
selected, deselected = [], []
for item in sorted(items, key=lambda i: i.nodeid):
(selected if hash(item.nodeid) % shards + 1 == shard else deselected).append(item)
items[:] = selected
config.hook.pytest_deselected(items=deselected)
# Note: set PYTHONHASHSEED=0 in CI so hash() is stable across processes
// mocha shard helper: pass --spec from a script
const glob = require('glob');
const shard = Number(process.env.SHARD ?? 1);
const shards = Number(process.env.SHARDS ?? 1);
const files = glob.sync('test/e2e/**/*.spec.js').sort();
const mine = files.filter((_, i) => i % shards === shard - 1);
console.log(mine.join(' '));
// package.json: "test:shard": "mocha $(node scripts/shard.js)"
// NUnit: filter by a category assigned from a stable hash of the test name
// Simplest reliable approach: split by test class into N .runsettings filters, or use
// dotnet test --filter "FullyQualifiedName~Shard1" with a naming convention.
// For dynamic sharding, compute the shard in a custom NUnit attribute:
public class ShardAttribute : NUnitAttribute, IApplyToTest
{
public void ApplyToTest(Test test)
{
int shard = int.Parse(Environment.GetEnvironmentVariable("SHARD") ?? "1");
int shards = int.Parse(Environment.GetEnvironmentVariable("SHARDS") ?? "1");
int bucket = Math.Abs(test.FullName.GetHashCode() % shards) + 1; // use a stable hash in practice
if (bucket != shard) test.RunState = RunState.Ignored;
}
}

Capturing Evidence on Failure

Every framework has a hook that runs after a failed test. Use it to save a screenshot, the page source, and the console errors collected through BiDi. The AI-Assisted Testing lesson has the code for all four languages; the CI side is just an artifact upload of that folder.

If you run through Grid with video enabled, fire a session event on failure so only failing videos are kept:

Keep videos for failed tests only (Grid 4.41+)
Selenium 4 Medium
// In your afterEach/@AfterMethod when the test failed:
((RemoteWebDriver) driver).fireSessionEvent("test:failed", Map.of("testName", testName));
// Grid video container started with SE_UPLOAD_FAILURE_SESSION_ONLY=true uploads only these
driver.fire_session_event("test:failed", {"testName": request.node.name})
await driver.fireSessionEvent('test:failed', { testName: this.currentTest.title });
((RemoteWebDriver)driver).FireSessionEvent("test:failed", new Dictionary<string, object> { ["testName"] = TestContext.CurrentContext.Test.Name });

CI-Only Flakiness and Its Causes

Symptom in CI onlyCauseFix
Elements “not visible” that are visible locallySmaller default window in headless--window-size=1366,768 or setSize
session deleted because of page crash/dev/shm too small in Docker--shm-size=2g or --disable-dev-shm-usage
Random timeouts under parallel loadToo many browsers per runnerReduce parallelism; one browser per 2 vCPU
Tests pass on retryTiming assumptions masked by a fast laptopFix waits; do not add retries
Different browser version than localRunner image updatedPin browserVersion via Selenium Manager
Fonts and layout differMissing fonts on LinuxInstall fonts-liberation, fonts-noto in the image or use Grid images
First test in each job slowCold Selenium Manager cacheCache SE_CACHE_PATH

Summary

  • Headless, pinned browsers, cached drivers, sized windows, explicit timeouts and artifacts on failure are the non-negotiables.
  • Run browsers in the official Selenium images as CI services when you need version control, video or more than one browser type.
  • Shard deterministically across jobs; match parallelism to runner resources.
  • Treat CI-only failures as real bugs in the tests’ timing assumptions, not as noise to retry away.

Copy-paste recipes for this topic

Related lessons