Teach Paige to identify document types
The three ways to add document types — preconfigured library, custom, or detected from samples — and the AI Document Type Identification description that decides whether each document gets classified as the right one.
At the end of this guide, your scan job has a set of document types the AI can reliably tell apart — because each one carries an identification description that names what makes it unmistakable.
The page says it plainly: the types under Your Active Document Types "will be used by AI and indexing to classify your scanned documents." Every document that comes through the job gets matched against this set — so the set, and how distinctly each type is described, decides how often classification is right.
First, the distinction that prevents most confusion
Two different AI decisions happen to every batch, configured on two different pages:
| Decision | Question it answers | Where it's configured |
|---|---|---|
| Separation | Where does one document end and the next begin? | Settings > scan job > Separation |
| Identification | What is each separated document? | Settings > scan job > Document Types |
If pages from two documents end up glued together, that's separation. If a cleanly separated document gets labeled as the wrong type, that's identification — this guide. Separation has its own AI guidance box ("guidance for the AI to help identify document boundaries"); don't put classification hints there, or identification hints in it.
Add document types
Open Settings (the gear at the bottom of the left menu — administrators only), pick the scan job, and choose Document Types. Three buttons, three strategies:
The Document Types page with its three add buttons and the Your Active Document Types section below
- Add From Preconfigured Library — a curated library of ready-made types, organized under Financial, Operations, Human Resources, and Legal & Compliance. Hover over a type to see its fields. These come with professionally written identification and extraction guidance already in place — the fastest path when your documents are standard forms like invoices or W-2s.
- Add Custom Type — build a type by hand: name, identification description, fields.
- Add from Sample Documents — upload sample documents, and "the system will automatically detect document types and fields. Each unique document will create its own document type." While it works you'll see a card reading Creating Document Type… / Analyzing uploaded document…. Review what it built afterward — auto-detected types are a starting draft, not a finished configuration.
Write the identification description
Open any type (the pencil icon, Edit Document Type). The second field on the form — right under Document Type Name — is AI Document Type Identification: information to help AI identify this type of document. The placeholder points at the right material: "unique features of this document type (e.g., logo position, specific text, layout)."
What separates a good description from a useless one is contrast with the other types in this job:
- Name the fixed text that always appears — a form number, a standard title, a legal phrase.
- Describe the layout — a table of line items, a boxed grid, a signature block at the bottom.
- Say what it is not, when two of your types look alike: if the job has both purchase orders and order confirmations, each description should name the words that settle it ("titled 'Purchase Order' with a PO number — not an order confirmation, which arrives on the vendor's letterhead").
The description is about recognizing the document. What to pull off of it belongs to each field's extraction guidance — a separate skill with its own guide.
Keep the set lean
Every type you add is another candidate the AI must consider for every document. Types you no longer receive are pure confusion risk — delete them (the trash icon; the confirmation means it: "This action is permanent."). A tight set of well-described types beats a long list of vague ones.
If classification still misses
The Advanced page of the scan job has a Document Type Model selector — the AI model used to identify document types, from Swift (fast and economical) through Efficient to Precision (advanced accuracy). If you've sharpened the descriptions and two similar types still get confused, stepping up the model is the remaining lever.
Success check: run a mixed test batch containing at least one document of each active type. Every document should come out labeled with the right type. Any miss tells you which two descriptions to sharpen against each other.
