AdemeroHelp

Train CapturePoint templates that keep working

Train a Document Layout the structured way — classification and field extraction as separate steps, validated with regular expressions, and proven against a batch of real samples before production.

Administrators45 minutesVerified 2026-08-30

Before you begin

  • Administrator access to the CapturePoint workstation, with configuration unlocked
  • The Worker Service running (CapturePoint won't process pages without it)
  • A stack of recent, real documents from each vendor you want to train

At the end of this guide you'll have a trained Document Layout that recognizes a vendor's paperwork and pulls the right values into the right fields — proven against a batch of real samples before production ever depends on it.

Note

CapturePoint calls these Document Layouts in the interface, not templates. Same idea: a trainable pattern for one vendor's stationery.

You're training two separate things

Most "the template broke" problems come from mixing these up:

QuestionWhat it's calledWhere it lives (navigation drawer)
"Which layout is this page?"ClassificationClassify > Classification Configuration / Classification Automation
"Where on this page is the value?"Field extractionIndex > Index-Field Configuration / Index-Field Automation

A page can classify perfectly and still extract nothing — and the reverse. When a field exports blank even though the text on the page is clean and readable, the culprit is almost always the layout's field blocks pointing at the wrong spot, not the OCR. Fix the layout, not the scanner.

Pick samples that teach one thing

  • One clean, straight, representative first page trains one layout. The wizard learns this stationery pattern — not the business meaning of every future document.
  • One layout = one stationery pattern. Two vendors, or one vendor's old and new invoice designs, are separate layouts (or separate layout images — see the last section).
  • Skewed, faxed, or marked-up pages make poor teachers. Save those for the test batch, where they belong.

Train a layout

  1. Capture your sample page, then open the navigation drawer and choose Begin Training Session. (If you don't see it, choose Unlock Configuration first.)

  1. When asked Would you like to train using the current image or all images?, choose CURRENT IMAGE — you're teaching from your one clean sample.

  1. When asked Does this image represent a document's first page or a supporting page?, choose FIRST PAGE for the page that starts each document. This answer drives how CapturePoint finds document boundaries later.

  1. At Select the recognition type, choose IMAGE SIGNATURE for ordinary stationery, or BARCODE if the page carries a reliable barcode — then drag and size a box over the barcode when prompted.

  1. For each field block the wizard suggests, pick one: Create new index field, Map to existing index field, or Ignore. Ignore anything you don't need — fewer fields means fewer things to break.

  1. Name the layout at Document Layout Name. Use vendor plus document type (for example Acme - Invoice) so the list stays readable at fifty layouts.

    What you should see: Training is complete.

  1. If the wizard offers Would you like to improve accuracy by adding another location? or Would you like to add another index rule?, take it — a second anchor location makes classification much harder to fool.

Note

If you retrain with a page the system already knows, it asks what to do with the existing layout. Group with Existing Document Layout adds your sample to it; only rename or recreate if you truly meant to replace it.

Figure

Training wizard showing suggested field blocks on a sample invoice, with the Create new index field / Map to existing index field / Ignore choices visible.

Fixed-location blocks: get the box right

Text and barcode rules use a Location Type of either Fixed Location or Presence on Page.

  • A Fixed Location is a box defined by Left, Top, Width, Height, and a Tolerance, all in inches. Draw the box a little generous — tolerance absorbs normal scanner drift, but a tight box around a value that moves will miss.
  • The properties dialog warns: there must be relevant text on the trained image at that location for the rule to work. If the box covers whitespace on your sample, the rule can never fire.
  • A Presence on Page rule searches the whole page instead, and requires a Regular Expression — it can't be blank.
  • For barcode rules, leave Barcode Type on Any unless misreads force you to lock it to one symbology (Code 39, Code 128, QR Code, and others are available).

Validate the value, not just the position

A block in the right place can still grab the wrong text. Add a Regular Expression to each extraction rule:

  1. In the rule's properties, pick a Regular Expression Type — prebuilt patterns include Numbers Only, Date with 4-Digit Year (12/25/2011), and Invoice #, Number Only (Invoice #: 1234) — or choose Custom and write your own.

  1. Type a real value from your sample into the Test Value box and watch the Test Result: line.

    What you should see: Success: True, Matches: 1. If it says Invalid Regular Expression, fix the pattern before saving.

  1. For critical fields, turn on Enable Confidence Scoring with an Alert Threshold, and mark must-have fields Required For Completion so weak reads stop for review instead of exporting quietly.

First page vs. supporting pages

Field extraction reads the first page of each document. In Document-Layout Properties (opened from the layout in the Classification Rules screen) you control that behavior: Enable First-Page Identifier, First Page is Cover Sheet, and Classify Only When First Page in Captured File.

Important

When checking a layout, always compare against your samples' first pages only. Judging field blocks against a page 2 tells you nothing — the rules never look there.

Prove it before production: a pass/fail exercise

Never promote a layout on the strength of its training sample. Test it:

  1. Gather at least 5 recent, real documents per layout — different dates, amounts, and scan quality. Capture them into a test job.

  1. Before touching anything, write down what you expect:

    Sample Invoice # Date Total Pass?
    1 10482 08/12/2026 1,240.00
    2 10517 08/19/2026 88.50
  1. On the Classification Rules screen, run Reprocess All Documents.

    What you should see: every sample classified to the correct layout, with the reason Classified using Document Layout.

  1. On the Index rules screen, click RECOGNIZE ALL and wait for Reprocessing Index Rules... to finish.

  1. Fill in the table from what actually extracted. The layout passes only when every field on every sample matches.

  • All samples classified to the correct layout — none unclassified, none stolen by another layout
  • Every expected field value extracted exactly, on every sample
  • No field-alert or empty-field surprises at export
Important

A blank field on a sample whose text is perfectly readable means the layout is wrong, not the OCR. Re-open the layout, check each field block's position against the failing sample's first page, and re-run RECOGNIZE ALL.

When a vendor changes their stationery

Layouts don't wear out — stationery changes under them. When a vendor's documents stop classifying or fields start drifting:

  1. Open the layout's properties from the Classification Rules screen and click MANAGE LAYOUT IMAGES.

  2. For a modest change, add a current sample to the layout — confirm Would you like to add this image to the current document layout?. The layout now recognizes both versions.

  3. For a full redesign, train it as a new layout instead (or use RECREATE LAYOUT to start that layout over).

  4. Re-run the pass/fail exercise above with fresh samples before trusting it again.

Still stuck?

If a layout still fails the pass/fail exercise after you've re-checked its field blocks against the failing samples' first pages, contact support with: the layout name, your filled-in pass/fail table (expected vs. extracted values), which samples classified wrong or extracted blank, and your CapturePoint version.