RAG/AI Test PDF Sets

Document-type-classified PDFs for testing chunking, embedding, and retrieval pipelines. Papers, slides, reports, contracts — each with predictable structure so you can validate segmentation behavior.

paper-style.pdf · 2-COL · FOOTNOTES · ZIP

01WHAT'S IN THE PACK

What's in the Pack

rag-test-pack.zip contains:

  • paper-style.pdfTwo-column academic layout, abstract, numbered sections, footnotes, references. Tests: heading-aware chunking, footnote handling, reference-list isolation.
  • slides-style.pdf15 slide-like pages, bullet points, large headers, minimal body text. Tests: short-chunk behavior, title extraction, sparse-content embedding.
  • report-style.pdfCorporate report with TOC, numbered headings, tables, charts-as-images. Tests: TOC extraction, nested-heading hierarchy, image-text separation.
  • contract-style.pdfNumbered clauses, defined terms, signature block, dense justified text. Tests: long-paragraph chunking, term extraction, section-boundary detection.
  • headers-footers-variant.pdfSame content as paper-style with repeating running headers and page numbers. Tests: header/footer noise removal before chunking.
  • README.mdPer-file documentation: structure summary, expected chunk boundaries, known gotchas.

02HOW TO USE

How to Use

  1. 01
    Download the ZIP and extract all files.
  2. 02
    Read README.md for per-file structure notes and expected chunking behavior.
  3. 03
    Ingest each PDF into your RAG pipeline.
  4. 04
    Compare actual chunks against the documented expected boundaries.
  5. 05
    Repeat with the header/footer variant to verify noise filtering.

03DOWNLOAD

Download

Download rag-test-pack.zip Generate a Custom RAG Test PDF → STATIC CDN · NO GENERATION DELAY · NO ACCOUNT

Four primary types — academic paper (two-column, footnotes), presentation slides (bullet-heavy, sparse), corporate report (TOC, tables, images), and legal contract (dense clauses, defined terms) — plus a header/footer noise variant.

No. All text is placeholder (lorem-ipsum-style) with realistic structure: real heading hierarchy, numbered sections, table layouts, and footnote markers. The structure is what matters for chunking tests, not the words.

The ZIP is the primary distribution format so the README stays with the files. For single custom files, use the generator at /generator.