Make already-captured documents text-searchable
The Catalog Manager backfill that OCRs a whole catalog of image documents — the options that matter, the support authorization it requires, and how to read its results, which report queuing, not completion.
Before you begin
- Access to the Content Central server (Catalog Manager runs there)
- An authorization code from Ademero support — request it when you're ready to run, it expires quickly
At the end of this guide, a catalog full of image-only documents — old scans, a bulk import, years of unsearchable TIFFs — is queued for OCR, and you'll know how to read the tool's output, because what it reports is queuing, and the conversion itself happens afterward, in the background.
First, confirm this is your problem: Why text you can see isn't searchable covers the diagnosis. This page is the bulk cure.
Run it
In Catalog Manager (on the server), select the catalog, choose Make Searchable selected catalog in the action drop-down at the bottom, and click Go. (Make Searchable all catalogs exists for the full sweep.)
The Catalog Make Searchable Warning dialog collects the decisions:
- Select the image types to make searchable: — PDF, TIFF, JPEG, Bitmap, PNG, GIF, all checked by default.
- Select the processing method:
- Convert to fully-searchable PDF (files may become large) — the documents become searchable PDFs.
- Make fully-searchable, but leave file in original form — the text goes into the search index; the files stay as they are. Choose this when the original format matters (or when "files may become large" gives you pause).
- Only process documents identified as not searchable — on by default, and usually right: it pre-scans the index so already-searchable documents aren't redone. (The pre-scan itself takes time on a large catalog — that's the pause before progress moves.)
- Ademero Support Authorization Code: — this operation is deliberately support-gated. Get the code from Ademero support when you're ready to run; it's time-limited, so same-day, not stockpiled.
OK, and let the progress dialog run to its summary.
Reading the results — the part that trips people
The summary reports, per catalog:
- "Documents newly added to the queue" — what this run queued.
- "Documents already waiting in the queue" — queued by an earlier run, not yet converted.
And then the sentence that explains the whole design, straight from the dialog: the queued documents are converted by the Capture Service — "if this number is not going down, check that the service is running and look in the Content Central log for entries beginning 'Make Searchable:'."
So: "finished" means queued. Conversion proceeds in the background at the server's pace; each converted document gets a new version (its history reads Make Searchable), and the search index picks it up from there. Two readings to know:
| You see | It means |
|---|---|
| Newly added: 0, already waiting: 5,000 | An earlier run's queue is still draining — or the Capture Service is stopped. Don't re-run; check the service |
| Newly added: 0, already waiting: 0 | There was nothing to do — everything matching your type selections is already searchable |
Success check: pick one document you know was image-only, wait for the queue to drain past it, and search for a phrase on its page — it should hit, and the document's history shows the Make Searchable row.
Keeping future captures searchable
The backfill shouldn't need a sequel. PDFs uploaded through the Capture page can prompt for a PDF action — Capture As Is, Make Fully Searchable, or Make Searchable (Keep Original) — and scanning paths OCR as they capture. If new unsearchable documents keep appearing, find the intake path that's skipping OCR rather than planning another backfill.
