# Getting Paige document separation right

Pick the right separation method for each scan job, write AI guidance that describes real boundary cues, and prove the setup on a small test set before trusting large batches.

Product: paige · Audience: system-administrator · Time: 30 minutes · Last verified: 2026-08-30

Canonical: https://help.ademero.com/paige/document-processing/paige-document-separation

**At the end of this guide your scan job will split incoming pages into the right documents — and you'll have a small, trusted test set that proves it before you run a large batch.**

## Files are not documents

One uploaded file is not necessarily one document. A single PDF from a scanning session can hold twenty invoices; a single emailed attachment can hold a contract plus unrelated correspondence. Paige separates pages into documents per scan job, and the method you choose changes everything downstream: document classification and field extraction both run per *document*, so a wrong split produces wrong types and wrong field values no matter how good the rest of your setup is.

| Method | Section | Use when |
| --- | --- | --- |
| **File Boundaries** | Simple | Every uploaded file is exactly one document |
| **Single Page** | Simple | Every document is a single page |
| **Barcode Separation** | Advanced | A barcode marks the start of each document |
| **AI Separation** | Advanced | Boundaries must be recognized from page content |

> **IMPORTANT:** **Single Page** and **File Boundaries** are exclusive modes. Turning either one on automatically turns off **Barcode Separation**, **AI Separation**, and the other simple mode. With **File Boundaries** active you'll see the banner "File Boundaries mode active: Each uploaded file will be treated as a separate document. No additional separation will be performed." — if your files contain multiple documents, this mode will merge them, by design.

> **NOTE:** For email scan jobs, the **Capture** page's Email Capture settings include a switch labeled **Allow document separation within each attachment**. It controls whether separation runs inside each emailed attachment.

## Open the separation settings

1. Open the **Settings** page. Your scan jobs are listed in the sidebar under **Scan Jobs**.

2. Select the scan job you're tuning, then choose **Separation** in the left menu.
   
   **What you should see:** a page starting with "Document separation methods determine how to identify individual documents," split into **Advanced Separation Methods** and **Simple Separation Methods**.

3. Turn on the method that matches your intake. For mixed multi-document files with no barcodes, that's **AI Separation** ("Let AI identify the start and stop of documents.").

## Write guidance that describes boundaries, not one example

With **AI Separation** on, click **Add AI Guidance** (or **Edit AI Guidance** once guidance exists) and enter your rules in the **Edit AI Guidance** dialog. Saved guidance appears on the card under **Current guidance:**.

Guidance that names one example document generalizes poorly. Guidance that describes what a *boundary between two adjacent pages* looks like works far better:

- Describe the first page of a new document: a fresh document number, a new date, letterhead or a title block reappearing, `Page 1 of N` numbering starting over.
- Describe continuation pages just as explicitly: running page numbers, "continued" wording, carried-forward totals, exhibit or schedule headings that belong to the document before them.
- Call out your hardest case directly. If vendors send packets with several invoices back to back, say that a new invoice number starts a new document *even when the vendor and layout are identical*.
- Say what is **not** a boundary: terms-and-conditions pages, attachments, and appendices should stay with their parent document.

## Build a small test set — with both kinds of mistakes covered

Paige has no built-in accuracy report for separation, so this is a working practice, not a product feature. Gather 15–30 real files and write down the right answer for each before you run anything:

- [ ] Files holding several documents each — including same-vendor multi-invoice packets (these catch **under-splitting**, where boundaries get missed)
- [ ] A few genuinely long single documents — contracts with exhibits, multi-page statements (these catch **over-splitting**, where false boundaries appear inside one document)
- [ ] A couple of ordinary one-document files as controls
- [ ] An answer key: for each file, the page number where each document starts

> **IMPORTANT:** Real batches usually contain both failure modes at once — some documents merged together *and* some long documents chopped apart. A test set that only covers one kind will approve a configuration that fails on the other.

## Validate on the small set before trusting large runs

1. Run the test set through the scan job and wait for the **Document Separation** stage to finish.

2. In the **Documents** area, open separation review — the toggle button's tooltip reads **Show Separation View**.

3. Compare each file's result against your answer key. Press `H` for the **Keyboard Shortcuts** dialog; to fix a missed boundary, select the page where the new document should start (**Shift + Click** selects a range, **Ctrl + Click** selects specific pages) and press **Enter** to add a document break. `Ctrl+Z` undoes a change.

4. Tally the errors in two columns — boundaries missed vs. false boundaries added. The two failure modes have different fixes: missed boundaries usually need sharper "what starts a document" cues; false boundaries usually need clearer "what a continuation page looks like" cues.
   
   **What you should see:** a near-perfect score on the small set. That's your green light for volume — if the small set isn't clean, a large run won't be either, and you'll be reviewing thousands of pages instead of dozens.

## Check your answer key before blaming the model

When Paige disagrees with your key, open the actual pages before counting it as an error. In practice, a meaningful share of "model mistakes" turn out to be labeling mistakes in the key — a configuration that looks mediocre can turn out to be scoring near-perfectly once the answer key is corrected. Fix the key, re-score, and only then decide whether the guidance needs another pass.

## If clean guidance still misses boundaries

Two settings people reach for — one helps, one is unrelated:

1. On the **Advanced** page, under **AI Models**, raise the **Document Separation Model** — "The AI model used to determine where one document ends and another begins." The default is `Swift - Fast and economical`; switch to `Precision - Advanced accuracy` for difficult batches. The choice is saved per scan job, so only your hard-mix jobs pay the extra cost.

2. Leave **Fields: First Page Only** (on the **Processing** page) out of it. That card controls whether field values are extracted from only the first page of each document — it affects field extraction, never where documents split.

## Still stuck?

If your answer key is verified, the guidance describes boundary cues, and the `Precision` model still gets a file wrong, contact Ademero support with this compact package:

- The exact file (the original upload, not screenshots)
- The expected boundaries: the page number where each document should start
- What Paige produced instead — which documents were merged, and which were split apart
- Your current AI guidance text and the selected **Document Separation Model**

## What's next

- [A Paige field came back blank or wrong](https://help.ademero.com/paige/document-processing/paige-extracted-blank-or-wrong-field)
