Skip to main content
SeleniumDecoded

Verify the Contents of a Downloaded File

Assert on a downloaded CSV, PDF or Excel file: read it after download completes, check headers and rows, and clean up.

Medium Verified against Selenium 4.48.0

Problem

A download completing is not the test. The test is that the export contains the right rows. Parsing the file needs a library per format and a small amount of care around encodings and temporary files.

Solution

Combine the download-wait recipe with a format-specific reader and assert on a few distinctive values rather than the whole file.

Read CSV and PDF exports
Selenium 4 Medium
// CSV with Apache Commons CSV (org.apache.commons:commons-csv)
try (Reader reader = Files.newBufferedReader(csvFile, StandardCharsets.UTF_8);
CSVParser parser = CSVFormat.DEFAULT.builder().setHeader().setSkipHeaderRecord(true).build().parse(reader)) {
List<CSVRecord> rows = parser.getRecords();
assertEquals(List.of("order_id", "customer", "total"), parser.getHeaderNames());
assertEquals(42, rows.size());
assertTrue(rows.stream().anyMatch(r -> r.get("order_id").equals("INV-1042") && r.get("total").equals("1240.00")));
}
// PDF with PDFBox (org.apache.pdfbox:pdfbox)
try (PDDocument doc = Loader.loadPDF(pdfFile.toFile())) {
String text = new PDFTextStripper().getText(doc);
assertTrue(text.contains("Invoice INV-1042"));
}
import csv
# CSV (stdlib)
with open(csv_file, newline="", encoding="utf-8-sig") as f: # utf-8-sig strips a BOM if Excel-style
rows = list(csv.DictReader(f))
assert list(rows[0].keys()) == ["order_id", "customer", "total"]
assert len(rows) == 42
assert any(r["order_id"] == "INV-1042" and r["total"] == "1240.00" for r in rows)
# PDF (pip install pypdf)
from pypdf import PdfReader
text = "".join(p.extract_text() for p in PdfReader(pdf_file).pages)
assert "Invoice INV-1042" in text
# Excel (pip install openpyxl)
from openpyxl import load_workbook
ws = load_workbook(xlsx_file, read_only=True).active
header = [c.value for c in next(ws.iter_rows(max_row=1))]
assert header == ["order_id", "customer", "total"]
// CSV: npm i csv-parse
const { parse } = require('csv-parse/sync');
const rows = parse(fs.readFileSync(csvFile, 'utf8'), { columns: true, bom: true });
assert.deepStrictEqual(Object.keys(rows[0]), ['order_id', 'customer', 'total']);
assert.strictEqual(rows.length, 42);
assert.ok(rows.some((r) => r.order_id === 'INV-1042' && r.total === '1240.00'));
// PDF: npm i pdf-parse
const pdfParse = require('pdf-parse');
const { text } = await pdfParse(fs.readFileSync(pdfFile));
assert.ok(text.includes('Invoice INV-1042'));
// CSV: NuGet CsvHelper
using var reader = new StreamReader(csvFile, Encoding.UTF8);
using var csv = new CsvReader(reader, CultureInfo.InvariantCulture);
csv.Read(); csv.ReadHeader();
Assert.That(csv.HeaderRecord, Is.EqualTo(new[] { "order_id", "customer", "total" }));
var rows = csv.GetRecords<dynamic>().ToList();
Assert.That(rows.Count, Is.EqualTo(42));
Assert.That(rows.Any(r => r.order_id == "INV-1042" && r.total == "1240.00"));
// PDF: NuGet PdfPig
using var doc = PdfDocument.Open(pdfFile);
var text = string.Join("", doc.GetPages().Select(p => p.Text));
Assert.That(text, Does.Contain("Invoice INV-1042"));

Why It Works

Reading the file after the completion wait guarantees a full, closed file. Asserting on headers, row count and one or two distinctive rows verifies the export logic without coupling to every value, which would make the test fail on any unrelated data change.

Gotchas

  • Excel-generated CSVs often start with a UTF-8 BOM; strip it (utf-8-sig, bom: true) or the first header becomes order_id.
  • PDF text extraction merges columns unpredictably; assert on substrings, not exact lines.
  • Delete the download directory in teardown; leftover files make the “exactly one file” wait ambiguous on the next test.
  • Through a Grid, fetch the file with managed downloads before reading; see the download-wait recipe.

Learn the theory