CI/CD Pipelines: GitHub Actions, GitLab CI and Jenkins
Run Selenium suites in CI reliably: headless browsers, Selenium Manager caching, Grid services in pipelines, parallel jobs, artifacts on failure, and the settings that stop CI-only flakiness.
A Selenium suite that only runs on a laptop is a demo. CI is where it earns its keep, and CI is also where most suites first turn flaky: different screen sizes, no GPU, slower disks, cold caches, and a browser version that changed overnight. This lesson gives you working pipeline definitions for the three most common CI systems and the settings that make them stable.
The Checklist for Any CI System
- Headless browser with an explicit window size.
- Pinned or cached browser and driver so runs are reproducible and fast. Selenium Manager handles the matching; you handle the cache.
- Parallelism matched to the runner’s CPU and memory (roughly one browser per 2 vCPUs and 2 GB).
- Artifacts on failure: screenshots, page source, console logs, videos if using Grid.
- Timeouts at every level: test, job, and driver, so a hung browser fails fast.
- Retry policy for known-flaky tags only, never blanket retries.
- Secrets for test accounts and cloud credentials from the CI secret store, never in the repository.
GitHub Actions
Ubuntu runners ship with Chrome, Firefox and Edge preinstalled, and Selenium Manager finds them. Cache its download directory so drivers are not fetched every run.
name: E2Eon: pull_request: schedule: - cron: '0 3 * * *' # nightly full run
jobs: selenium: runs-on: ubuntu-latest timeout-minutes: 30 strategy: fail-fast: false matrix: browser: [chrome, firefox] shard: [1, 2, 3] env: BROWSER: ${{ matrix.browser }} HEADLESS: 'true' BASE_URL: https://staging.example.com SE_CACHE_PATH: ${{ github.workspace }}/.se-cache steps: - uses: actions/checkout@v4
- uses: actions/setup-java@v4 with: { distribution: temurin, java-version: '21', cache: maven }
- name: Cache Selenium Manager downloads uses: actions/cache@v4 with: path: .se-cache key: se-${{ runner.os }}-${{ matrix.browser }}-${{ hashFiles('pom.xml') }}
- name: Run shard ${{ matrix.shard }} run: mvn -B test -Dshard=${{ matrix.shard }} -Dshards=3 env: TEST_USER: ${{ secrets.TEST_USER }} TEST_PASSWORD: ${{ secrets.TEST_PASSWORD }}
- name: Upload failure evidence if: failure() uses: actions/upload-artifact@v4 with: name: evidence-${{ matrix.browser }}-${{ matrix.shard }} path: | target/screenshots/ target/surefire-reports/ retention-days: 7The same shape works for Python (actions/setup-python and pytest -n 2 --dist loadfile), Node (actions/setup-node and npx mocha --parallel) and .NET (actions/setup-dotnet and dotnet test).
Running Against a Grid Service Container
When you need a specific browser version or want video, run the official Selenium image as a service and point GRID_URL at it.
services: selenium: image: selenium/standalone-chrome:4.48.0 ports: ['4444:4444'] options: >- --shm-size=2g --health-cmd "curl -sf http://localhost:4444/status || exit 1" --health-interval 5s --health-timeout 5s --health-retries 20 env: GRID_URL: http://localhost:4444--shm-size=2g matters: Chrome crashes with “session deleted because of page crash” when /dev/shm is the default 64 MB.
GitLab CI
GitLab’s services keyword works the same way. The job container needs your language runtime; the browser lives in the service.
e2e: stage: test image: python:3.12 services: - name: selenium/standalone-chrome:4.48.0 alias: selenium variables: SE_NODE_MAX_SESSIONS: "2" SE_VNC_NO_PASSWORD: "1" variables: GRID_URL: http://selenium:4444 BASE_URL: https://staging.example.com HEADLESS: "true" FF_NETWORK_PER_BUILD: "true" # services and job share a network cache: key: pip-$CI_COMMIT_REF_SLUG paths: [.cache/pip] before_script: - pip install -r requirements.txt - | for i in $(seq 1 30); do curl -sf http://selenium:4444/status | grep -q '"ready": true' && break sleep 2 done script: - pytest tests/e2e -n 2 --junitxml=report.xml artifacts: when: always reports: junit: report.xml paths: [failures/] expire_in: 1 week parallel: 3 # GitLab sets CI_NODE_INDEX / CI_NODE_TOTAL for sharding timeout: 30 minutesSharding with parallel: 3 requires your test runner to split by CI_NODE_INDEX; pytest can do it with the pytest-split plugin, or partition by file name hash in a conftest.py hook.
Jenkins
A declarative pipeline running tests against a Docker-hosted Grid, with the Grid started in a sidecar and torn down in post.
// Jenkinsfilepipeline { agent { label 'docker' } options { timeout(time: 40, unit: 'MINUTES'); timestamps() } environment { GRID_URL = 'http://localhost:4444' BASE_URL = 'https://staging.example.com' HEADLESS = 'true' TEST_CREDS = credentials('e2e-test-user') // exposes TEST_CREDS_USR and TEST_CREDS_PSW } stages { stage('Grid up') { steps { sh ''' docker run -d --name grid --shm-size=2g -p 4444:4444 \ -e SE_NODE_MAX_SESSIONS=4 selenium/standalone-chrome:4.48.0 for i in $(seq 1 30); do curl -sf http://localhost:4444/status | grep -q '"ready": true' && exit 0 sleep 2 done exit 1 ''' } } stage('Test') { steps { sh 'mvn -B test -Dparallel.threads=4' } } } post { always { junit 'target/surefire-reports/*.xml' archiveArtifacts artifacts: 'target/screenshots/**', allowEmptyArchive: true sh 'docker logs grid > grid.log 2>&1 || true; docker rm -f grid || true' archiveArtifacts artifacts: 'grid.log', allowEmptyArchive: true } }}Sharding Tests Across Jobs
Parallel jobs need a deterministic way to split tests. A stable approach in any language: sort test identifiers, then take every Nth.
// JUnit 5: a condition that skips tests not in this shardpublic class ShardCondition implements ExecutionCondition { @Override public ConditionEvaluationResult evaluateExecutionCondition(ExtensionContext ctx) { int shard = Integer.getInteger("shard", 1); int shards = Integer.getInteger("shards", 1); String id = ctx.getUniqueId(); int bucket = Math.floorMod(id.hashCode(), shards) + 1; return bucket == shard ? ConditionEvaluationResult.enabled("in shard " + shard) : ConditionEvaluationResult.disabled("not in shard " + shard); }}// Register globally via META-INF/services/org.junit.jupiter.api.extension.Extension// with -Djunit.jupiter.extensions.autodetection.enabled=trueimport os
def pytest_collection_modifyitems(config, items): shard = int(os.getenv("CI_NODE_INDEX", "1")) shards = int(os.getenv("CI_NODE_TOTAL", "1")) if shards == 1: return selected, deselected = [], [] for item in sorted(items, key=lambda i: i.nodeid): (selected if hash(item.nodeid) % shards + 1 == shard else deselected).append(item) items[:] = selected config.hook.pytest_deselected(items=deselected)
# Note: set PYTHONHASHSEED=0 in CI so hash() is stable across processes// mocha shard helper: pass --spec from a scriptconst glob = require('glob');const shard = Number(process.env.SHARD ?? 1);const shards = Number(process.env.SHARDS ?? 1);
const files = glob.sync('test/e2e/**/*.spec.js').sort();const mine = files.filter((_, i) => i % shards === shard - 1);console.log(mine.join(' '));// package.json: "test:shard": "mocha $(node scripts/shard.js)"// NUnit: filter by a category assigned from a stable hash of the test name// Simplest reliable approach: split by test class into N .runsettings filters, or use// dotnet test --filter "FullyQualifiedName~Shard1" with a naming convention.// For dynamic sharding, compute the shard in a custom NUnit attribute:public class ShardAttribute : NUnitAttribute, IApplyToTest{ public void ApplyToTest(Test test) { int shard = int.Parse(Environment.GetEnvironmentVariable("SHARD") ?? "1"); int shards = int.Parse(Environment.GetEnvironmentVariable("SHARDS") ?? "1"); int bucket = Math.Abs(test.FullName.GetHashCode() % shards) + 1; // use a stable hash in practice if (bucket != shard) test.RunState = RunState.Ignored; }}Capturing Evidence on Failure
Every framework has a hook that runs after a failed test. Use it to save a screenshot, the page source, and the console errors collected through BiDi. The AI-Assisted Testing lesson has the code for all four languages; the CI side is just an artifact upload of that folder.
If you run through Grid with video enabled, fire a session event on failure so only failing videos are kept:
// In your afterEach/@AfterMethod when the test failed:((RemoteWebDriver) driver).fireSessionEvent("test:failed", Map.of("testName", testName));// Grid video container started with SE_UPLOAD_FAILURE_SESSION_ONLY=true uploads only thesedriver.fire_session_event("test:failed", {"testName": request.node.name})await driver.fireSessionEvent('test:failed', { testName: this.currentTest.title });((RemoteWebDriver)driver).FireSessionEvent("test:failed", new Dictionary<string, object> { ["testName"] = TestContext.CurrentContext.Test.Name });CI-Only Flakiness and Its Causes
| Symptom in CI only | Cause | Fix |
|---|---|---|
| Elements “not visible” that are visible locally | Smaller default window in headless | --window-size=1366,768 or setSize |
session deleted because of page crash | /dev/shm too small in Docker | --shm-size=2g or --disable-dev-shm-usage |
| Random timeouts under parallel load | Too many browsers per runner | Reduce parallelism; one browser per 2 vCPU |
| Tests pass on retry | Timing assumptions masked by a fast laptop | Fix waits; do not add retries |
| Different browser version than local | Runner image updated | Pin browserVersion via Selenium Manager |
| Fonts and layout differ | Missing fonts on Linux | Install fonts-liberation, fonts-noto in the image or use Grid images |
| First test in each job slow | Cold Selenium Manager cache | Cache SE_CACHE_PATH |
Summary
- Headless, pinned browsers, cached drivers, sized windows, explicit timeouts and artifacts on failure are the non-negotiables.
- Run browsers in the official Selenium images as CI services when you need version control, video or more than one browser type.
- Shard deterministically across jobs; match parallelism to runner resources.
- Treat CI-only failures as real bugs in the tests’ timing assumptions, not as noise to retry away.
Copy-paste recipes for this topic
- Pick the Browser, Headless Mode and Grid From Environment Variables
Run the same suite on Chrome, Firefox or Edge, locally or on a Grid, headed or headless, by changing environment variables instead of code.