Verify the Contents of a Downloaded File
Assert on a downloaded CSV, PDF or Excel file: read it after download completes, check headers and rows, and clean up.
Medium Verified against Selenium 4.48.0
Problem
A download completing is not the test. The test is that the export contains the right rows. Parsing the file needs a library per format and a small amount of care around encodings and temporary files.
Solution
Combine the download-wait recipe with a format-specific reader and assert on a few distinctive values rather than the whole file.
Read CSV and PDF exports
Selenium 4 Medium
// CSV with Apache Commons CSV (org.apache.commons:commons-csv)try (Reader reader = Files.newBufferedReader(csvFile, StandardCharsets.UTF_8); CSVParser parser = CSVFormat.DEFAULT.builder().setHeader().setSkipHeaderRecord(true).build().parse(reader)) { List<CSVRecord> rows = parser.getRecords(); assertEquals(List.of("order_id", "customer", "total"), parser.getHeaderNames()); assertEquals(42, rows.size()); assertTrue(rows.stream().anyMatch(r -> r.get("order_id").equals("INV-1042") && r.get("total").equals("1240.00")));}
// PDF with PDFBox (org.apache.pdfbox:pdfbox)try (PDDocument doc = Loader.loadPDF(pdfFile.toFile())) { String text = new PDFTextStripper().getText(doc); assertTrue(text.contains("Invoice INV-1042"));}import csv
# CSV (stdlib)with open(csv_file, newline="", encoding="utf-8-sig") as f: # utf-8-sig strips a BOM if Excel-style rows = list(csv.DictReader(f))assert list(rows[0].keys()) == ["order_id", "customer", "total"]assert len(rows) == 42assert any(r["order_id"] == "INV-1042" and r["total"] == "1240.00" for r in rows)
# PDF (pip install pypdf)from pypdf import PdfReadertext = "".join(p.extract_text() for p in PdfReader(pdf_file).pages)assert "Invoice INV-1042" in text
# Excel (pip install openpyxl)from openpyxl import load_workbookws = load_workbook(xlsx_file, read_only=True).activeheader = [c.value for c in next(ws.iter_rows(max_row=1))]assert header == ["order_id", "customer", "total"]// CSV: npm i csv-parseconst { parse } = require('csv-parse/sync');const rows = parse(fs.readFileSync(csvFile, 'utf8'), { columns: true, bom: true });assert.deepStrictEqual(Object.keys(rows[0]), ['order_id', 'customer', 'total']);assert.strictEqual(rows.length, 42);assert.ok(rows.some((r) => r.order_id === 'INV-1042' && r.total === '1240.00'));
// PDF: npm i pdf-parseconst pdfParse = require('pdf-parse');const { text } = await pdfParse(fs.readFileSync(pdfFile));assert.ok(text.includes('Invoice INV-1042'));// CSV: NuGet CsvHelperusing var reader = new StreamReader(csvFile, Encoding.UTF8);using var csv = new CsvReader(reader, CultureInfo.InvariantCulture);csv.Read(); csv.ReadHeader();Assert.That(csv.HeaderRecord, Is.EqualTo(new[] { "order_id", "customer", "total" }));var rows = csv.GetRecords<dynamic>().ToList();Assert.That(rows.Count, Is.EqualTo(42));Assert.That(rows.Any(r => r.order_id == "INV-1042" && r.total == "1240.00"));
// PDF: NuGet PdfPigusing var doc = PdfDocument.Open(pdfFile);var text = string.Join("", doc.GetPages().Select(p => p.Text));Assert.That(text, Does.Contain("Invoice INV-1042"));Why It Works
Reading the file after the completion wait guarantees a full, closed file. Asserting on headers, row count and one or two distinctive rows verifies the export logic without coupling to every value, which would make the test fail on any unrelated data change.
Gotchas
- Excel-generated CSVs often start with a UTF-8 BOM; strip it (
utf-8-sig,bom: true) or the first header becomesorder_id. - PDF text extraction merges columns unpredictably; assert on substrings, not exact lines.
- Delete the download directory in teardown; leftover files make the “exactly one file” wait ambiguous on the next test.
- Through a Grid, fetch the file with managed downloads before reading; see the download-wait recipe.