# Train CapturePoint templates that keep working

Train a Document Layout the structured way — classification and field extraction as separate steps, validated with regular expressions, and proven against a batch of real samples before production.

Product: capturepoint · Audience: system-administrator · Time: 45 minutes · Last verified: 2026-08-30

Canonical: https://help.ademero.com/capturepoint/templates/train-capturepoint-templates

**At the end of this guide you'll have a trained Document Layout that recognizes a vendor's paperwork and pulls the right values into the right fields — proven against a batch of real samples before production ever depends on it.**

> **NOTE:** CapturePoint calls these **Document Layouts** in the interface, not templates. Same idea: a trainable pattern for one vendor's stationery.

## You're training two separate things

Most "the template broke" problems come from mixing these up:

| Question | What it's called | Where it lives (navigation drawer) |
| --- | --- | --- |
| "Which layout is this page?" | Classification | **Classify** > **Classification Configuration** / **Classification Automation** |
| "Where on this page is the value?" | Field extraction | **Index** > **Index-Field Configuration** / **Index-Field Automation** |

A page can classify perfectly and still extract nothing — and the reverse. When a field exports blank even though the text on the page is clean and readable, the culprit is almost always the layout's field blocks pointing at the wrong spot, **not** the OCR. Fix the layout, not the scanner.

## Pick samples that teach one thing

- One clean, straight, representative **first page** trains one layout. The wizard learns *this stationery pattern* — not the business meaning of every future document.
- **One layout = one stationery pattern.** Two vendors, or one vendor's old and new invoice designs, are separate layouts (or separate layout images — see the last section).
- Skewed, faxed, or marked-up pages make poor teachers. Save those for the test batch, where they belong.

## Train a layout

1. Capture your sample page, then open the navigation drawer and choose **Begin Training Session**. (If you don't see it, choose **Unlock Configuration** first.)

2. When asked `Would you like to train using the current image or all images?`, choose **CURRENT IMAGE** — you're teaching from your one clean sample.

3. When asked `Does this image represent a document's first page or a supporting page?`, choose **FIRST PAGE** for the page that starts each document. This answer drives how CapturePoint finds document boundaries later.

4. At **Select the recognition type**, choose **IMAGE SIGNATURE** for ordinary stationery, or **BARCODE** if the page carries a reliable barcode — then drag and size a box over the barcode when prompted.

5. For each field block the wizard suggests, pick one: **Create new index field**, **Map to existing index field**, or **Ignore**. Ignore anything you don't need — fewer fields means fewer things to break.

6. Name the layout at **Document Layout Name**. Use vendor plus document type (for example `Acme - Invoice`) so the list stays readable at fifty layouts.
   
   **What you should see:** `Training is complete.`

7. If the wizard offers `Would you like to improve accuracy by adding another location?` or `Would you like to add another index rule?`, take it — a second anchor location makes classification much harder to fool.

> **NOTE:** If you retrain with a page the system already knows, it asks what to do with the existing layout. **Group with Existing Document Layout** adds your sample to it; only rename or recreate if you truly meant to replace it.

*[Screenshot: Training wizard showing suggested field blocks on a sample invoice, with the Create new index field / Map to existing index field / Ignore choices visible.]*

## Fixed-location blocks: get the box right

Text and barcode rules use a **Location Type** of either **Fixed Location** or **Presence on Page**.

- A **Fixed Location** is a box defined by **Left**, **Top**, **Width**, **Height**, and a **Tolerance**, all in inches. Draw the box a little generous — tolerance absorbs normal scanner drift, but a tight box around a value that moves will miss.
- The properties dialog warns: there must be relevant text on the trained image at that location for the rule to work. If the box covers whitespace on your sample, the rule can never fire.
- A **Presence on Page** rule searches the whole page instead, and requires a **Regular Expression** — it can't be blank.
- For barcode rules, leave **Barcode Type** on `Any` unless misreads force you to lock it to one symbology (`Code 39`, `Code 128`, `QR Code`, and others are available).

## Validate the value, not just the position

A block in the right place can still grab the wrong text. Add a **Regular Expression** to each extraction rule:

1. In the rule's properties, pick a **Regular Expression Type** — prebuilt patterns include `Numbers Only`, `Date with 4-Digit Year (12/25/2011)`, and `Invoice #, Number Only (Invoice #: 1234)` — or choose `Custom` and write your own.

2. Type a real value from your sample into the **Test Value** box and watch the **Test Result:** line.
   
   **What you should see:** `Success: True, Matches: 1`. If it says `Invalid Regular Expression`, fix the pattern before saving.

3. For critical fields, turn on **Enable Confidence Scoring** with an **Alert Threshold**, and mark must-have fields **Required For Completion** so weak reads stop for review instead of exporting quietly.

## First page vs. supporting pages

Field extraction reads the **first page** of each document. In **Document-Layout Properties** (opened from the layout in the Classification Rules screen) you control that behavior: **Enable First-Page Identifier**, **First Page is Cover Sheet**, and **Classify Only When First Page in Captured File**.

> **IMPORTANT:** When checking a layout, always compare against your samples' **first pages only**. Judging field blocks against a page 2 tells you nothing — the rules never look there.

## Prove it before production: a pass/fail exercise

Never promote a layout on the strength of its training sample. Test it:

1. Gather at least 5 recent, real documents per layout — different dates, amounts, and scan quality. Capture them into a test job.

2. Before touching anything, write down what you expect:
   
   | Sample | Invoice # | Date | Total | Pass? |
   |---|---|---|---|---|
   | 1 | `10482` | `08/12/2026` | `1,240.00` | |
   | 2 | `10517` | `08/19/2026` | `88.50` | |

3. On the **Classification Rules** screen, run **Reprocess All Documents**.
   
   **What you should see:** every sample classified to the correct layout, with the reason `Classified using Document Layout`.

4. On the Index rules screen, click **RECOGNIZE ALL** and wait for `Reprocessing Index Rules...` to finish.

5. Fill in the table from what actually extracted. The layout passes only when **every field on every sample** matches.

- [ ] All samples classified to the correct layout — none unclassified, none stolen by another layout
- [ ] Every expected field value extracted exactly, on every sample
- [ ] No field-alert or empty-field surprises at export

> **IMPORTANT:** A blank field on a sample whose text is perfectly readable means the layout is wrong, not the OCR. Re-open the layout, check each field block's position against the failing sample's first page, and re-run **RECOGNIZE ALL**.

## When a vendor changes their stationery

Layouts don't wear out — stationery changes under them. When a vendor's documents stop classifying or fields start drifting:

1. Open the layout's properties from the Classification Rules screen and click **MANAGE LAYOUT IMAGES**.
2. For a modest change, add a current sample to the layout — confirm `Would you like to add this image to the current document layout?`. The layout now recognizes both versions.
3. For a full redesign, train it as a new layout instead (or use **RECREATE LAYOUT** to start that layout over).
4. Re-run the pass/fail exercise above with fresh samples before trusting it again.

## Still stuck?

If a layout still fails the pass/fail exercise after you've re-checked its field blocks against the failing samples' first pages, contact support with: the layout name, your filled-in pass/fail table (expected vs. extracted values), which samples classified wrong or extracted blank, and your CapturePoint version.

## What's next

- [A field extracted blank or wrong in CapturePoint](https://help.ademero.com/capturepoint/templates/capturepoint-extracted-blank-or-wrong-field)
- [Tell CapturePoint where documents start and end with QCards](https://help.ademero.com/capturepoint/templates/qcards-where-documents-start-and-end)
