01WHAT'S IN THE PACK
What's in the Pack
ocr-test-pack.zip contains:
- clean-sans-serif.pdfHigh-contrast, single-column, 12pt sans-serif text. No ligatures, no kerning anomalies. Purpose: accuracy baseline.
- tables.pdfMulti-column data tables with merged cells, numeric alignment, and thin rules. Purpose: table-structure recognition and cell-boundary detection.
- two-column.pdfAcademic two-column layout with justified text and inline math symbols. Purpose: reading-order detection, column segmentation.
- low-quality-scan.pdfSimulated 150 DPI scan: slight rotation (0.5°), uneven margins, toner-fade artifacts, JPEG compression noise. Purpose: robustness under degraded input.
- ground-truth.txtUTF-8 text file with the exact expected content of each PDF, delimited by filename headers. Use for CER/WER computation.
- README.mdPer-file notes: what each variant tests, recommended evaluation metrics.
02HOW TO USE
How to Use
- 01Download and extract the ZIP.
- 02Run your OCR engine on each PDF.
- 03Compare output against the corresponding section in ground-truth.txt.
- 04Compute Character Error Rate (CER) or Word Error Rate (WER).
- 05Use the clean-sans-serif baseline to verify your pipeline works before testing harder variants.
03DOWNLOAD
Download
Download ocr-test-pack.zip
Generate a Scan-Style PDF →
STATIC CDN · NO GENERATION DELAY · NO ACCOUNT
Character Error Rate (CER) and Word Error Rate (WER) against the ground-truth text. The clean-sans-serif file serves as an upper-bound baseline; degraded variants reveal where your engine loses accuracy.
It is a digitally rendered PDF that simulates common scan artifacts: slight rotation, uneven margins, toner fade, and JPEG compression. It is deterministic and reproducible, unlike a real scan.